Should you use gpt-oss-120b?
gpt-oss-120b is listed for chat, coding workloads and supports a 131K context window.
Use it when Cerebras's free tier is enough for evaluation, demos, or light production traffic.
https://api.cerebras.ai/v1 gpt-oss-120b gpt-oss-120b is listed for chat, coding workloads and supports a 131K context window.
Use it when Cerebras's free tier is enough for evaluation, demos, or light production traffic.
We only found this gpt-oss-120b listing on Cerebras in the current catalog. Use the related models section for nearby alternatives.
| Provider | Model listing | Access | Context | API | Limits |
|---|---|---|---|---|---|
| | gpt-oss-120b | Free tier | 131K | OpenAI-style | 5 RPM, 30K TPM, 1M TPD |
| Model | Provider | Context | Access |
|---|---|---|---|
| Llama 3.1 70B | Cerebras | 131K | Free tier |
| zai-glm-4.7 (deprecated Aug 2026) | Cerebras | 131K | Free tier |
| gemma-4-31b | Cerebras | 131K | Free tier |
| zai-glm-4.7 | Cerebras | 128K | Free tier |
| llama-3.3-70b | Cerebras | 128K | Check provider |
gpt-oss-120b is tagged for chat in this catalog and works with OpenAI-compatible client libraries.
gpt-oss-120b is tagged for coding in this catalog and works with OpenAI-compatible client libraries.
gpt-oss-120b is listed with free API access on Cerebras, subject to the provider's quota and account policy.
The model ID shown in this catalog is gpt-oss-120b.
The listed free tier limit is 5 RPM, 30K TPM, 1M TPD. Limits can change per account tier, so confirm against the provider dashboard.
The listed context window is 131K tokens with up to 32K output tokens.
Cerebras offers GPT-oss-120b, a powerful text-based LLM ideal for chat and coding, generating up to 8,000 tokens from 128,000-token contexts at 30 RPM, 14,400 RPD, and 1M TPD, without requiring a credit card and compatible with OpenAI.
For API keys, setup steps, and provider-level limits, see the Cerebras provider page.