Cerebras logo

Free Cerebras API Key, Base URL & Rate Limits

API Provider

Cerebras Cloud offers free API access to Llama and GPT-OSS models running on the Cerebras Wafer-Scale Engine, one of the fastest AI accelerators available.

OpenAI CompatibleFree TierFunction CallingVision

How to get a free Cerebras API key

  1. 1
    Sign up at cloud.cerebras.ai Email or GitHub. No credit card.
  2. 2
    Go to API Keys
  3. 3
    Generate an API key
  4. 4
    Choose a model Llama 3.3 70B or GPT-OSS 120B available for free.
  5. 5
    Configure OpenAI client Base URL: https://api.cerebras.ai/v1

Provider Snapshot

P Provider Type API Provider
O OpenAI Compatible Yes
U Base URL https://api.cerebras.ai/v1
F Free Tier 5 free models online
P Phone Required No
S Streaming Varies
T Function Calling Yes
V Vision Yes
L Last Updated 2026-08-06

Supported Models 8 models

View in directory →
Model ID Developer Context Availability Free Tier Use cases
gemma-4-31b
gemma-4-31b
Google 131K Online Yes chat
zai-glm-4.7
zai-glm-4.7
Cerebras 128K Online Yes chat
zai-glm-4.7
zai-glm-4.7 (deprecated Aug 2026)
Cerebras 131K Online Yes chat
gpt-oss-120b
gpt-oss-120b
OpenAI 131K Online Yes chatcoding
llama3.1-70b
Llama 3.1 70B
Meta 131K Online Yes chatcoding
qwen-3-235b-a22b-instruct-2507
qwen-3-235b-a22b-instruct-2507
Alibaba 131K Check provider Yes chat
qwen-3-32b
qwen-3-32b
Alibaba 131K Check provider Yes chat
llama-3-3-70b
llama-3.3-70b
Meta 128K Check provider Yes chat

Developer Tools

K Save API Key Never lose your API keys — save keys from all providers in one place. T Test API Key Test your API key and check quota.

Cerebras FreeLLM Score free API access score

How we score →
71 /100
✅ Solid Choice — Strong in easy signup
Generosity Free limits and access terms
85
Access Signup and key availability
100
Model Breadth Free model coverage
50
Reliability Status and stability signals
65
Compatibility SDK and endpoint support
70
Quality Model capability signals
55

What is Cerebras?

Ultra-fast inference on Cerebras WSE chips — 1M tokens/day.

Cerebras Cloud offers free API access to Llama and GPT-OSS models running on the Cerebras Wafer-Scale Engine, one of the fastest AI accelerators available. The free tier provides 1 million tokens/day and 14,400 requests/day per model with no credit card required. Context window is limited to 8K on the free tier.

  • Ultra-fast inference on WSE chips
  • 1M tokens/day free
  • No credit card required
  • Llama 3.1 8B + GPT-OSS 120B available

API Compatibility: Live-tested OpenAI-compatible Chat Completions

Cerebras Free Tier Limits & Pricing

Credit Card Not required
Phone Verification Not required
Free Tier 5 free models online
Context Range 128K – 131K
Total Models 8 listed
Rate Limits 15 RPM, 30K TPM, 1M TPD · 10 RPM, 100 RPD, 1M TPD · 5 RPM, 30K TPM, 1M TPD
API Compatibility Live-tested OpenAI-compatible Chat Completions

Cerebras API Setup Tutorials

Cerebras is fully compatible with popular AI coding assistants like Cursor, Claude Code, and more. To see step-by-step API configuration instructions for your favorite tool, please visit our Global Configuration Guide →

Cerebras Model Use Cases

What Cerebras's free models are best for, based on aggregated model capabilities:

Chat 8 models Coding 2 models

Cerebras Limitations & Caveats

  • 8K context window on free tier (vs 128K on paid)
  • Limited model selection — Llama and GPT-OSS only
  • 1M tokens/day shared across models

Cerebras FAQ

Why is Cerebras limited to 8K context on the free tier?

Cerebras' WSE chips are optimized for throughput, not context length. The free tier caps context at 8K tokens to ensure fair resource allocation. Paid tier unlocks 128K context.

What is GPT-OSS 120B on Cerebras?

GPT-OSS (Open Source Software) is a 120B parameter model trained by Cerebras and made available through their platform. It's a general-purpose model comparable to GPT-4 class models for many tasks.

Is Cerebras as fast as Groq?

Both are hardware-accelerated (Cerebras WSE vs Groq LPU). In benchmarks, Cerebras achieves similar throughput to Groq on comparable models — typically 1,500-2,500 tokens/second.