Should you use zai-glm-4.7?
zai-glm-4.7 is listed for chat workloads and supports a 128K context window.
Use it when Cerebras's free tier is enough for evaluation, demos, or light production traffic.
GLM-4.7 is Z.AI's (Zhipu AI) latest-generation bilingual model, offered free through Cerebras Cloud on its WSE inference hardware.
https://api.cerebras.ai/v1 zai-glm-4.7 zai-glm-4.7 is listed for chat workloads and supports a 128K context window.
Use it when Cerebras's free tier is enough for evaluation, demos, or light production traffic.
We only found this glm listing on Cerebras in the current catalog. Use the related models section for nearby alternatives.
| Provider | Model listing | Access | Context | API | Limits |
|---|---|---|---|---|---|
| | zai-glm-4.7 | Free tier | 128K | OpenAI-style | 10 RPM, 100 RPD, 1M TPD |
| Model | Provider | Context | Access |
|---|---|---|---|
| Llama 3.1 70B | Cerebras | 131K | Free tier |
| gpt-oss-120b | Cerebras | 131K | Free tier |
| zai-glm-4.7 (deprecated Aug 2026) | Cerebras | 131K | Free tier |
| gemma-4-31b | Cerebras | 131K | Free tier |
| llama-3.3-70b | Cerebras | 128K | Check provider |
zai-glm-4.7 is tagged for chat in this catalog and works with OpenAI-compatible client libraries.
zai-glm-4.7 is listed with free API access on Cerebras, subject to the provider's quota and account policy.
The model ID shown in this catalog is zai-glm-4.7.
The listed free tier limit is 10 RPM, 100 RPD, 1M TPD. Limits can change per account tier, so confirm against the provider dashboard.
The listed context window is 128K tokens with up to 8K output tokens.
GLM-4.7 is Z.AI's (Zhipu AI) latest-generation bilingual model, offered free through Cerebras Cloud on its WSE inference hardware. With 128K context, OpenAI-compatible API, and no credit card requirement, it is a strong choice for developers who need Chinese-English bilingual performance or want an alternative to the Llama/Qwen families. The free tier is more constrained than some Cerebras endpoints — 100 requests per day at 10 RPM with a 1M token daily cap — so it is best used for evaluation, comparison testing, or low-volume bilingual tasks rather than continuous production traffic.
For API keys, setup steps, and provider-level limits, see the Cerebras provider page.