Should you use GLM-4.5-Flash?
GLM-4.5-Flash is listed for chat workloads and supports a 128K context window.
Use it when Z AI (Zhipu AI)'s free tier is enough for evaluation, demos, or light production traffic.
https://open.bigmodel.cn/api/paas/v4 glm-4.5 | Benchmark | Score | Metric | Date |
|---|---|---|---|
| Artificial Analysis Coding Index | 26.3 | index | 2026-03-11 |
| SciCode | 34.8 | percent correct | 2026-03-11 |
| Terminal-Bench Hard | 22 | success rate | 2026-03-11 |
GLM-4.5-Flash is listed for chat workloads and supports a 128K context window.
Use it when Z AI (Zhipu AI)'s free tier is enough for evaluation, demos, or light production traffic.
We only found this glm-4-5 listing on Z AI (Zhipu AI) in the current catalog. Use the related models section for nearby alternatives.
| Provider | Model listing | Access | Context | API | Limits |
|---|---|---|---|---|---|
| | GLM-4.5-Flash | Free tier | 128K | OpenAI-style | 1 concurrent request |
| Model | Provider | Context | Access |
|---|---|---|---|
| GLM-4.7-Flash | Z AI (Zhipu AI) | 200K | Free tier |
| GLM-4.6V-Flash | Z AI (Zhipu AI) | 128K | Free tier |
| GLM-4.5-Air | Z AI (Zhipu AI) | 131K | Free tier |
GLM-4.5-Flash is tagged for chat in this catalog and works with OpenAI-compatible client libraries.
GLM-4.5-Flash is listed with free API access on Z AI (Zhipu AI), subject to the provider's quota and account policy.
The model ID shown in this catalog is glm-4.5.
The listed free tier limit is 1 concurrent request. Limits can change per account tier, so confirm against the provider dashboard.
The listed context window is 128K tokens with up to 96K output tokens.
Z AI's GLM-4.5-Flash offers a free LLM for text tasks, generating up to 8,000 tokens of context from 128,000 tokens input, with no credit card needed and OpenAI compatibility, ideal for chat applications with a rate limit of one concurrent request.
For API keys, setup steps, and provider-level limits, see the Z AI (Zhipu AI) provider page.