Should you use llama-3.3-70b-versatile?
llama-3.3-70b-versatile is listed for chat workloads and supports a 131K context window.
Use the comparison and availability sections before choosing this listing, because it is not currently marked as free.
https://api.groq.com/openai/v1 llama-3.3-70b-versatile | Benchmark | Score | Metric | Date |
|---|---|---|---|
| Artificial Analysis Coding Index | 10.7 | index | 2026-03-11 |
| SciCode | 26 | percent correct | 2026-03-11 |
| Terminal-Bench Hard | 3 | success rate | 2026-03-11 |
llama-3.3-70b-versatile is listed for chat workloads and supports a 131K context window.
Use the comparison and availability sections before choosing this listing, because it is not currently marked as free.
We found 5 provider listings for llama-3-3-70b-instruct. Check model ID, quota, pricing, and API format before switching providers.
| Provider | Model listing | Access | Context | API | Limits |
|---|---|---|---|---|---|
| | llama-3.3-70b-versatile | Paid | 131K | OpenAI-style | 30 RPM, 1,000 RPD |
| | @cf/meta/llama-3.3-70b-instruct-fp8-fast | Free tier | 131K | Native | 10K neurons/day (shared) |
| | Llama-3.3-70B-Instruct | Free tier | 131K | OpenAI-style | 15 RPM, 150 RPD |
| | Meta-Llama-3_3-70B-Instruct | Free tier | 131K | OpenAI-style | 2 RPM (anonymous) |
| | Llama-3.3-70B-Instruct | Free tier | 128K | Native | Varies |
| Model | Provider | Context | Access |
|---|---|---|---|
| @cf/meta/llama-3.3-70b-instruct-fp8-fast | Cloudflare Workers AI | 131K | Free tier |
| Llama-3.3-70B-Instruct | GitHub Models | 131K | Free tier |
| Meta-Llama-3_3-70B-Instruct | OVHcloud AI Endpoints | 131K | Free tier |
| Llama-3.3-70B-Instruct | NVIDIA NIM | 128K | Free tier |
| Moonshot Kimi K2 | Groq | 131K | Free tier |
llama-3.3-70b-versatile is tagged for chat in this catalog and works with OpenAI-compatible client libraries.
llama-3.3-70b-versatile is not currently marked as free on Groq.
The model ID shown in this catalog is llama-3.3-70b-versatile.
The listed free tier limit is 30 RPM, 1,000 RPD. Limits can change per account tier, so confirm against the provider dashboard.
The listed context window is 131K tokens with up to 32K output tokens.
Llama 3.3 70B on Groq delivers Meta's flagship 70B model with Groq's ultra-fast LPU inference — expect dramatically lower latency compared to GPU-based providers. With 131K context, 32K output, and OpenAI SDK compatibility, it is one of the fastest ways to access a proven 70B-class model for interactive applications. The free tier is generous: 14,400 requests per day at 30 RPM, making it viable for moderate production workloads. Registration is required but no credit card is needed. If your application values response speed above all else, this Groq + Llama 3.3 combination is hard to beat.
For API keys, setup steps, and provider-level limits, see the Groq provider page.