Should you use gpt-oss:120b?
gpt-oss:120b is listed for chat, coding workloads and supports a 128K context window.
Use it when Ollama Cloud's free tier is enough for evaluation, demos, or light production traffic.
https://ollama.com/api gpt-oss:120b gpt-oss:120b is listed for chat, coding workloads and supports a 128K context window.
Use it when Ollama Cloud's free tier is enough for evaluation, demos, or light production traffic.
We only found this gpt-oss-120b listing on Ollama Cloud in the current catalog. Use the related models section for nearby alternatives.
| Provider | Model listing | Access | Context | API | Limits |
|---|---|---|---|---|---|
| | gpt-oss:120b | Free tier | 128K | OpenAI-style | Session/weekly limits (unpublished) |
| Model | Provider | Context | Access |
|---|---|---|---|
| deepseek-v4-pro | Ollama Cloud | 1.0M | Check provider |
| deepseek-v4-flash | Ollama Cloud | 1.0M | Check provider |
| minimax-m3 | Ollama Cloud | 512K | Free tier |
| nemotron-3-ultra | Ollama Cloud | 262K | Free tier |
| gpt-oss:20b | Ollama Cloud | 131K | Free tier |
gpt-oss:120b is tagged for chat in this catalog and works with OpenAI-compatible client libraries.
gpt-oss:120b is tagged for coding in this catalog and works with OpenAI-compatible client libraries.
gpt-oss:120b is listed with free API access on Ollama Cloud, subject to the provider's quota and account policy.
The model ID shown in this catalog is gpt-oss:120b.
The listed free tier limit is Session/weekly limits (unpublished). Limits can change per account tier, so confirm against the provider dashboard.
The listed context window is 128K tokens with up to 131K output tokens.
Cerebras offers GPT-oss-120b, a powerful text-based LLM ideal for chat and coding, generating up to 8,000 tokens from 128,000-token contexts at 30 RPM, 14,400 RPD, and 1M TPD, without requiring a credit card and compatible with OpenAI.
For API keys, setup steps, and provider-level limits, see the Ollama Cloud provider page.