Groq logo

llama-3.3-70b-versatile pricing and alternatives

Paid / limited
★★★★★★★★★★ 3.5 Benchmark-backed score

Popular open Llama workhorse for multilingual chat, coding, and self-hosting

Paid / limitedOpenAI compatibleTool callingJSON modeText
API Details

Copy-ready connection details

Use these values in your client or SDK.
Base URL
https://api.groq.com/openai/v1
Model ID
llama-3.3-70b-versatile
API format OpenAI Chat Completions + OpenAI Responses
No current free listing. This page is kept for comparison and SEO continuity. Check the alternatives below before integrating.
Technical Details

llama-3.3-70b-versatile specifications

Provider and model catalog
Context window 131K
Max output 32K
Status Online
Family llama
Knowledge cutoff 2023-12
Released Dec 6, 2024
Last updated Aug 6, 2026
Free listing since Dec 6, 2024
Input text
Output text
Capabilities tool calling, structured output, file attachments, temperature control
Open weights Yes

External benchmark references

Benchmark Score Metric Date
Artificial Analysis Coding Index 10.7 index 2026-03-11
SciCode 26 percent correct 2026-03-11
Terminal-Bench Hard 3 success rate 2026-03-11
AI Recommendation

Should you use llama-3.3-70b-versatile?

llama-3.3-70b-versatile is listed for chat workloads and supports a 131K context window.

Use the comparison and availability sections before choosing this listing, because it is not currently marked as free.

Best for
  • Chat
Strengths & Weaknesses

Strengths

  • Long context window
  • Tool calling support
  • Structured JSON output
  • Open weights available

Watch outs

  • No current free listing on this provider
  • Free-tier rate limits apply
  • Vision support is not listed
Benchmark Overview

Benchmark signals for llama-3.3-70b-versatile

Measured data
Intelligence General reasoning and instruction following
9.4/100
Coding Programming and code generation
11.9/100
Agentic Tool use and multi-step tasks
0.3/100
Speed Observed generation speed
80 tok/s
Context Maximum listed context window
131K
Pricing

llama-3.3-70b-versatile pricing per 1M tokens

Provider pricing
Input $0.6 per 1M tokens
Output $0.72 per 1M tokens
Free access Not listed Groq
Rate limit 30 RPM, 1,000 RPD provider policy
Availability

llama-3.3-70b-versatile availability by provider

4 alternatives

We found 5 provider listings for llama-3-3-70b-instruct. Check model ID, quota, pricing, and API format before switching providers.

Provider Model listing Access Context API Limits
Groq llama-3.3-70b-versatile Paid 131K OpenAI-style 30 RPM, 1,000 RPD
Cloudflare Workers AI @cf/meta/llama-3.3-70b-instruct-fp8-fast Free tier 131K Native 10K neurons/day (shared)
GitHub Models Llama-3.3-70B-Instruct Free tier 131K OpenAI-style 15 RPM, 150 RPD
OVHcloud AI Endpoints Meta-Llama-3_3-70B-Instruct Free tier 131K OpenAI-style 2 RPM (anonymous)
NVIDIA NIM Llama-3.3-70B-Instruct Free tier 128K Native Varies
View Groq setup guide →
Typical Use Cases

llama-3.3-70b-versatile use cases

Chat

llama-3.3-70b-versatile is tagged for chat in this catalog and works with OpenAI-compatible client libraries.

FAQ

llama-3.3-70b-versatile free API FAQ

Is llama-3.3-70b-versatile free to use?

llama-3.3-70b-versatile is not currently marked as free on Groq.

What is the llama-3.3-70b-versatile model ID?

The model ID shown in this catalog is llama-3.3-70b-versatile.

What are the llama-3.3-70b-versatile free tier rate limits on Groq?

The listed free tier limit is 30 RPM, 1,000 RPD. Limits can change per account tier, so confirm against the provider dashboard.

What context window does llama-3.3-70b-versatile support?

The listed context window is 131K tokens with up to 32K output tokens.

More about llama-3.3-70b-versatile

Llama 3.3 70B on Groq delivers Meta's flagship 70B model with Groq's ultra-fast LPU inference — expect dramatically lower latency compared to GPU-based providers. With 131K context, 32K output, and OpenAI SDK compatibility, it is one of the fastest ways to access a proven 70B-class model for interactive applications. The free tier is generous: 14,400 requests per day at 30 RPM, making it viable for moderate production workloads. Registration is required but no credit card is needed. If your application values response speed above all else, this Groq + Llama 3.3 combination is hard to beat.

For API keys, setup steps, and provider-level limits, see the Groq provider page.