Should you use Llama-4-Scout-17B-16E-Instruct?
Llama-4-Scout-17B-16E-Instruct is listed for chat workloads and supports a 512K context window.
Use it when GitHub Models's free tier is enough for evaluation, demos, or light production traffic.
Llama 4 Scout 17B on Groq runs Meta's latest MoE generation model with Groq's ultra-fast LPU inference.
https://models.github.ai/inference llama-4-scout-17b-16e-instruct Llama-4-Scout-17B-16E-Instruct is listed for chat workloads and supports a 512K context window.
Use it when GitHub Models's free tier is enough for evaluation, demos, or light production traffic.
We only found this llama listing on GitHub Models in the current catalog. Use the related models section for nearby alternatives.
| Provider | Model listing | Access | Context | API | Limits |
|---|---|---|---|---|---|
| | Llama-4-Scout-17B-16E-Instruct | Free tier | 512K | OpenAI-style | 15 RPM, 150 RPD |
| Model | Provider | Context | Access |
|---|---|---|---|
| Phi-4 | GitHub Models | 131K | Free tier |
| Mistral Large (24.11) | GitHub Models | 131K | Free tier |
| AI21 Jamba 1.5 Large | GitHub Models | 256K | Free tier |
| gpt-5 | GitHub Models | 200K | Free tier |
| gpt-4.1 | GitHub Models | 1.0M | Free tier |
Llama-4-Scout-17B-16E-Instruct is tagged for chat in this catalog and works with OpenAI-compatible client libraries.
Llama-4-Scout-17B-16E-Instruct is listed with free API access on GitHub Models, subject to the provider's quota and account policy.
The model ID shown in this catalog is llama-4-scout-17b-16e-instruct.
The listed free tier limit is 15 RPM, 150 RPD. Limits can change per account tier, so confirm against the provider dashboard.
The listed context window is 512K tokens with up to 4K output tokens.
Llama 4 Scout 17B on Groq runs Meta's latest MoE generation model with Groq's ultra-fast LPU inference. The Scout variant uses 16 active experts to deliver broad capability in a compact 17B active footprint, with 8K output per request. Combined with Groq's sub-200ms time-to-first-token, it offers a responsive experience for interactive chat and agent workflows. Rate limits are 14,400 requests per day at 30 RPM — sufficient for sustained prototyping and light production use. OpenAI SDK compatible; registration required but no credit card needed.
For API keys, setup steps, and provider-level limits, see the GitHub Models provider page.