Should you use gemma-4-31b?
gemma-4-31b is listed for chat workloads and supports a 131K context window.
Use it when Cerebras's free tier is enough for evaluation, demos, or light production traffic.
https://api.cerebras.ai/v1 gemma-4-31b gemma-4-31b is listed for chat workloads and supports a 131K context window.
Use it when Cerebras's free tier is enough for evaluation, demos, or light production traffic.
We found 2 provider listings for gemma. Check model ID, quota, pricing, and API format before switching providers.
| Provider | Model listing | Access | Context | API | Limits |
|---|---|---|---|---|---|
| | gemma-4-31b | Free tier | 131K | OpenAI-style | 15 RPM, 30K TPM, 1M TPD |
| | gemma-4-31B-it (Preview) | Free tier | 128K | Native | 20 RPM, 20 RPD, 200K TPD |
| Model | Provider | Context | Access |
|---|---|---|---|
| gemma-4-31B-it (Preview) | SambaNova | 128K | Free tier |
| Llama 3.1 70B | Cerebras | 131K | Free tier |
| gpt-oss-120b | Cerebras | 131K | Free tier |
| zai-glm-4.7 (deprecated Aug 2026) | Cerebras | 131K | Check provider |
| zai-glm-4.7 | Cerebras | 128K | Check provider |
gemma-4-31b is tagged for chat in this catalog and works with OpenAI-compatible client libraries.
gemma-4-31b is listed with free API access on Cerebras, subject to the provider's quota and account policy.
The model ID shown in this catalog is gemma-4-31b.
The listed free tier limit is 15 RPM, 30K TPM, 1M TPD. Limits can change per account tier, so confirm against the provider dashboard.
The listed context window is 131K tokens with up to 32K output tokens.
Largest Gemma 4 instruction model for open, self-hosted chat and reasoning
For API keys, setup steps, and provider-level limits, see the Cerebras provider page.