Should you use Llama-3.3-70B-Instruct?
Llama-3.3-70B-Instruct is listed for chat workloads and supports a 131K context window.
Use it when GitHub Models's free tier is enough for evaluation, demos, or light production traffic.
Popular open Llama workhorse for multilingual chat, coding, and self-hosting
https://models.github.ai/inference llama-3-3-70b-instruct | Benchmark | Score | Metric | Date |
|---|---|---|---|
| Artificial Analysis Coding Index | 10.7 | index | 2026-03-11 |
| SciCode | 26 | percent correct | 2026-03-11 |
| Terminal-Bench Hard | 3 | success rate | 2026-03-11 |
Llama-3.3-70B-Instruct is listed for chat workloads and supports a 131K context window.
Use it when GitHub Models's free tier is enough for evaluation, demos, or light production traffic.
We found 4 provider listings for llama-3-3-70b-instruct. Check model ID, quota, pricing, and API format before switching providers.
| Provider | Model listing | Access | Context | API | Limits |
|---|---|---|---|---|---|
| | Llama-3.3-70B-Instruct | Free tier | 131K | OpenAI-style | 15 RPM, 150 RPD |
| | @cf/meta/llama-3.3-70b-instruct-fp8-fast | Free tier | 131K | Native | 10K neurons/day (shared) |
| | Meta-Llama-3_3-70B-Instruct | Free tier | 131K | OpenAI-style | 2 RPM (anonymous) |
| | Llama-3.3-70B-Instruct | Free tier | 128K | Native | Varies |
| Model | Provider | Context | Access |
|---|---|---|---|
| @cf/meta/llama-3.3-70b-instruct-fp8-fast | Cloudflare Workers AI | 131K | Free tier |
| Meta-Llama-3_3-70B-Instruct | OVHcloud AI Endpoints | 131K | Free tier |
| Llama-3.3-70B-Instruct | NVIDIA NIM | 128K | Free tier |
| Phi-4 | GitHub Models | 131K | Free tier |
| Mistral Large (24.11) | GitHub Models | 131K | Free tier |
Llama-3.3-70B-Instruct is tagged for chat in this catalog and works with OpenAI-compatible client libraries.
Llama-3.3-70B-Instruct is listed with free API access on GitHub Models, subject to the provider's quota and account policy.
The model ID shown in this catalog is llama-3-3-70b-instruct.
The listed free tier limit is 15 RPM, 150 RPD. Limits can change per account tier, so confirm against the provider dashboard.
The listed context window is 131K tokens with up to 4K output tokens.
Llama-3.3-70B-Instruct — free model from GitHub Models.
For API keys, setup steps, and provider-level limits, see the GitHub Models provider page.