Z AI (Zhipu AI) logo

GLM-4.6V-Flash Free API on Z AI (Zhipu AI)

Free API Verified
★★★★★★★★★★ 3.5 Benchmark-backed score

Late GLM-4 workhorse for coding agents, reasoning, and structured tasks

Free APIOpenAI compatibleReasoningTool callingJSON modeTextImage
API Details

Copy-ready connection details

Use these values in your client or SDK.
Base URL
https://open.bigmodel.cn/api/paas/v4
Model ID
glm-4.6
API format OpenAI Chat Completions
Technical Details

GLM-4.6V-Flash specifications

Provider and model catalog
Context window 128K
Max output 4K
Status Online
Family glm
Knowledge cutoff 2025-04
Released Sep 30, 2025
Last updated Aug 6, 2026
Free listing since Dec 8, 2025
Input text
Output text
Capabilities reasoning, tool calling, structured output, temperature control
Open weights Yes

External benchmark references

Benchmark Score Metric Date
Artificial Analysis Coding Index 29.5 index 2026-05-22
SciCode 38.4 percent correct 2026-05-22
Terminal-Bench Hard 25 success rate 2026-05-22
SWE-Bench Pro 9.67 resolve rate Not listed
AI Recommendation

Should you use GLM-4.6V-Flash?

GLM-4.6V-Flash is listed for chat workloads and supports a 128K context window.

Use it when Z AI (Zhipu AI)'s free tier is enough for evaluation, demos, or light production traffic.

Best for
  • Chat
Strengths & Weaknesses

Strengths

  • Long context window
  • Reasoning mode listed
  • Tool calling support
  • Structured JSON output

Watch outs

  • Free-tier rate limits apply
Benchmark Overview

Benchmark signals for GLM-4.6V-Flash

Measured data
Intelligence General reasoning and instruction following
28.7/100
Coding Programming and code generation
45.8/100
Agentic Tool use and multi-step tasks
17.7/100
Context Maximum listed context window
128K
Pricing

GLM-4.6V-Flash pricing per 1M tokens

Free tier listed
Input $0.6 per 1M tokens
Output $2.2 per 1M tokens
Free access Available Z AI (Zhipu AI)
Rate limit 1 concurrent request provider policy
Availability

GLM-4.6V-Flash availability by provider

Current provider only

We only found this glm-4-6 listing on Z AI (Zhipu AI) in the current catalog. Use the related models section for nearby alternatives.

Provider Model listing Access Context API Limits
Z AI (Zhipu AI) GLM-4.6V-Flash Free tier 128K OpenAI-style 1 concurrent request
View Z AI (Zhipu AI) setup guide →
Typical Use Cases

GLM-4.6V-Flash use cases

Chat

GLM-4.6V-Flash is tagged for chat in this catalog and works with OpenAI-compatible client libraries.

FAQ

GLM-4.6V-Flash free API FAQ

Is GLM-4.6V-Flash free to use?

GLM-4.6V-Flash is listed with free API access on Z AI (Zhipu AI), subject to the provider's quota and account policy.

What is the GLM-4.6V-Flash model ID?

The model ID shown in this catalog is glm-4.6.

What are the GLM-4.6V-Flash free tier rate limits on Z AI (Zhipu AI)?

The listed free tier limit is 1 concurrent request. Limits can change per account tier, so confirm against the provider dashboard.

What context window does GLM-4.6V-Flash support?

The listed context window is 128K tokens with up to 4K output tokens.

More about GLM-4.6V-Flash

GLM-4.6V-Flash is a Z AI multimodal flash model listed with a 128,000 token context window and up to 4,000 output tokens. This entry tracks its OpenAI-compatible access, no-credit-card signal, and conservative one-request-at-a-time rate limit for developers comparing low-friction GLM options.

For API keys, setup steps, and provider-level limits, see the Z AI (Zhipu AI) provider page.