Free Cerebras API Key, Base URL & Rate Limits
API ProviderCerebras Cloud offers free API access to Llama and GPT-OSS models running on the Cerebras Wafer-Scale Engine, one of the fastest AI accelerators available.
How to get a free Cerebras API key
- 1
- 2 Go to API Keys
- 3 Generate an API key
- 4 Choose a model Llama 3.3 70B or GPT-OSS 120B available for free.
- 5 Configure OpenAI client Base URL: https://api.cerebras.ai/v1
Provider Snapshot
Supported Models 8 models
View in directory →| Model ID | Developer | Context | Availability | Free Tier | Use cases |
|---|---|---|---|---|---|
| gemma-4-31b gemma-4-31b | 131K | Online | Yes | chat | |
| zai-glm-4.7 zai-glm-4.7 | Cerebras | 128K | Online | Yes | chat |
| zai-glm-4.7 zai-glm-4.7 (deprecated Aug 2026) | Cerebras | 131K | Online | Yes | chat |
| gpt-oss-120b gpt-oss-120b | OpenAI | 131K | Online | Yes | chatcoding |
| llama3.1-70b Llama 3.1 70B | Meta | 131K | Online | Yes | chatcoding |
| qwen-3-235b-a22b-instruct-2507 qwen-3-235b-a22b-instruct-2507 | Alibaba | 131K | Check provider | Yes | chat |
| qwen-3-32b qwen-3-32b | Alibaba | 131K | Check provider | Yes | chat |
| llama-3-3-70b llama-3.3-70b | Meta | 128K | Check provider | Yes | chat |
Developer Tools
Cerebras FreeLLM Score free API access score
How we score →What is Cerebras?
Ultra-fast inference on Cerebras WSE chips — 1M tokens/day.
Cerebras Cloud offers free API access to Llama and GPT-OSS models running on the Cerebras Wafer-Scale Engine, one of the fastest AI accelerators available. The free tier provides 1 million tokens/day and 14,400 requests/day per model with no credit card required. Context window is limited to 8K on the free tier.
- Ultra-fast inference on WSE chips
- 1M tokens/day free
- No credit card required
- Llama 3.1 8B + GPT-OSS 120B available
API Compatibility: Live-tested OpenAI-compatible Chat Completions
Cerebras Free Tier Limits & Pricing
Cerebras API Setup Tutorials
Cerebras is fully compatible with popular AI coding assistants like Cursor, Claude Code, and more. To see step-by-step API configuration instructions for your favorite tool, please visit our Global Configuration Guide →
Cerebras Model Use Cases
What Cerebras's free models are best for, based on aggregated model capabilities:
Cerebras Limitations & Caveats
- 8K context window on free tier (vs 128K on paid)
- Limited model selection — Llama and GPT-OSS only
- 1M tokens/day shared across models
Cerebras FAQ
Why is Cerebras limited to 8K context on the free tier?
Cerebras' WSE chips are optimized for throughput, not context length. The free tier caps context at 8K tokens to ensure fair resource allocation. Paid tier unlocks 128K context.
What is GPT-OSS 120B on Cerebras?
GPT-OSS (Open Source Software) is a 120B parameter model trained by Cerebras and made available through their platform. It's a general-purpose model comparable to GPT-4 class models for many tasks.
Is Cerebras as fast as Groq?
Both are hardware-accelerated (Cerebras WSE vs Groq LPU). In benchmarks, Cerebras achieves similar throughput to Groq on comparable models — typically 1,500-2,500 tokens/second.