Cerebras logo

llama-3.3-70b API status on Cerebras

Check provider
★★★★★★★★★★ 3.5 Benchmark-backed score

Popular open Llama workhorse for multilingual chat, coding, and self-hosting

Check providerOpenAI compatibleTool callingJSON modeText
API Details

Copy-ready connection details

Use these values in your client or SDK.
Base URL
https://api.cerebras.ai/v1
Model ID
llama-3-3-70b
API format OpenAI Chat Completions
Technical Details

llama-3.3-70b specifications

Provider and model catalog
Context window 128K
Max output 8K
Status Check provider
Family llama
Knowledge cutoff 2023-12
Released Dec 6, 2024
Last updated Jun 15, 2026
Free listing since Dec 6, 2024
Input text
Output text
Capabilities tool calling, structured output, file attachments, temperature control
Open weights Yes

External benchmark references

Benchmark Score Metric Date
Artificial Analysis Coding Index 10.7 index 2026-03-11
SciCode 26 percent correct 2026-03-11
Terminal-Bench Hard 3 success rate 2026-03-11
AI Recommendation

Should you use llama-3.3-70b?

llama-3.3-70b is listed for chat workloads and supports a 128K context window.

Check Cerebras's current endpoint status before integrating this listing. Use the alternatives table if the provider listing is unavailable.

Best for
  • Chat
Strengths & Weaknesses

Strengths

  • Long context window
  • Tool calling support
  • Structured JSON output
  • Open weights available

Watch outs

  • Current endpoint availability needs provider confirmation
  • Free-tier rate limits apply
  • Vision support is not listed
Benchmark Overview

Benchmark signals for llama-3.3-70b

Measured data
Intelligence General reasoning and instruction following
9.4/100
Coding Programming and code generation
11.9/100
Agentic Tool use and multi-step tasks
0.3/100
Speed Observed generation speed
91 tok/s
Context Maximum listed context window
128K
Pricing

llama-3.3-70b pricing per 1M tokens

Check provider
Input $0.58 per 1M tokens
Output $0.71 per 1M tokens
Free access Check provider Cerebras
Rate limit 30 RPM, 14,400 RPD, 1M TPD provider policy
Availability

llama-3.3-70b availability by provider

4 alternatives

We found 5 provider listings for llama-3-3-70b-instruct. Check model ID, quota, pricing, and API format before switching providers.

Provider Model listing Access Context API Limits
Cerebras llama-3.3-70b Check provider 128K OpenAI-style 30 RPM, 14,400 RPD, 1M TPD
Cloudflare Workers AI @cf/meta/llama-3.3-70b-instruct-fp8-fast Free tier 131K Native 10K neurons/day (shared)
GitHub Models Llama-3.3-70B-Instruct Free tier 131K OpenAI-style 15 RPM, 150 RPD
OVHcloud AI Endpoints Meta-Llama-3_3-70B-Instruct Free tier 131K OpenAI-style 2 RPM (anonymous)
NVIDIA NIM Llama-3.3-70B-Instruct Free tier 128K Native Varies
View Cerebras setup guide →
Typical Use Cases

llama-3.3-70b use cases

Chat

llama-3.3-70b is tagged for chat in this catalog and works with OpenAI-compatible client libraries.

FAQ

llama-3.3-70b free API FAQ

Is llama-3.3-70b free to use?

llama-3.3-70b appears in the free model catalog for Cerebras, but its current endpoint availability should be confirmed with the provider before use.

What is the llama-3.3-70b model ID?

The model ID shown in this catalog is llama-3-3-70b.

What are the llama-3.3-70b free tier rate limits on Cerebras?

The listed free tier limit is 30 RPM, 14,400 RPD, 1M TPD. Limits can change per account tier, so confirm against the provider dashboard.

What context window does llama-3.3-70b support?

The listed context window is 128K tokens with up to 8K output tokens.

More about llama-3.3-70b

llama-3.3-70b — free model from Cerebras.

For API keys, setup steps, and provider-level limits, see the Cerebras provider page.