GitHub Models logo

Llama-4-Scout-17B-16E-Instruct Free API on GitHub Models

Free API
Catalog profile Tool use provider catalog metadata

Llama 4 Scout 17B on Groq runs Meta's latest MoE generation model with Groq's ultra-fast LPU inference.

Free APIOpenAI compatibleTool callingTextImage
API Details

Copy-ready connection details

Use these values in your client or SDK.
Base URL
https://models.github.ai/inference
Model ID
llama-4-scout-17b-16e-instruct
API format OpenAI Chat Completions
Technical Details

Llama-4-Scout-17B-16E-Instruct specifications

Provider catalog
Context window 512K
Max output 4K
Status Online
Family llama
Knowledge cutoff 2024-08
Released Apr 5, 2025
Last updated Aug 5, 2026
Free listing since Apr 5, 2025
Input text, image
Output text
Capabilities tool calling
Open weights Yes
AI Recommendation

Should you use Llama-4-Scout-17B-16E-Instruct?

Llama-4-Scout-17B-16E-Instruct is listed for chat workloads and supports a 512K context window.

Use it when GitHub Models's free tier is enough for evaluation, demos, or light production traffic.

Best for
  • Chat
Strengths & Weaknesses

Strengths

  • Long context window
  • Tool calling support
  • Works with OpenAI-style SDKs

Watch outs

  • Free-tier rate limits apply
Pricing

Llama-4-Scout-17B-16E-Instruct pricing per 1M tokens

Free tier listed
Input $0 per 1M tokens
Output $0 per 1M tokens
Free access Available GitHub Models
Rate limit 15 RPM, 150 RPD provider policy
Availability

Llama-4-Scout-17B-16E-Instruct availability by provider

Current provider only

We only found this llama listing on GitHub Models in the current catalog. Use the related models section for nearby alternatives.

Provider Model listing Access Context API Limits
GitHub Models Llama-4-Scout-17B-16E-Instruct Free tier 512K OpenAI-style 15 RPM, 150 RPD
View GitHub Models setup guide →
Typical Use Cases

Llama-4-Scout-17B-16E-Instruct use cases

Chat

Llama-4-Scout-17B-16E-Instruct is tagged for chat in this catalog and works with OpenAI-compatible client libraries.

FAQ

Llama-4-Scout-17B-16E-Instruct free API FAQ

Is Llama-4-Scout-17B-16E-Instruct free to use?

Llama-4-Scout-17B-16E-Instruct is listed with free API access on GitHub Models, subject to the provider's quota and account policy.

What is the Llama-4-Scout-17B-16E-Instruct model ID?

The model ID shown in this catalog is llama-4-scout-17b-16e-instruct.

What are the Llama-4-Scout-17B-16E-Instruct free tier rate limits on GitHub Models?

The listed free tier limit is 15 RPM, 150 RPD. Limits can change per account tier, so confirm against the provider dashboard.

What context window does Llama-4-Scout-17B-16E-Instruct support?

The listed context window is 512K tokens with up to 4K output tokens.

More about Llama-4-Scout-17B-16E-Instruct

Llama 4 Scout 17B on Groq runs Meta's latest MoE generation model with Groq's ultra-fast LPU inference. The Scout variant uses 16 active experts to deliver broad capability in a compact 17B active footprint, with 8K output per request. Combined with Groq's sub-200ms time-to-first-token, it offers a responsive experience for interactive chat and agent workflows. Rate limits are 14,400 requests per day at 30 RPM — sufficient for sustained prototyping and light production use. OpenAI SDK compatible; registration required but no credit card needed.

For API keys, setup steps, and provider-level limits, see the GitHub Models provider page.