NVIDIA NIM logo

glm-5.3-flash Free API on NVIDIA NIM

Free API Verified
★★★★★★★★★★ 3.5 Benchmark-backed score

z-ai/glm-5.3-flash — free model from NVIDIA NIM (z-ai).

Free APIOpenAI compatibleReasoningTool callingTextImageVideo
API Details

Copy-ready connection details

Use these values in your client or SDK.
Base URL
https://integrate.api.nvidia.com/v1
Model ID
z-ai/glm-5.3-flash
API format OpenAI Chat Completions
Technical Details

glm-5.3-flash specifications

Provider catalog
Context window 1.3M
Max output 131K
Status Online
Family glm
Released Aug 26, 2026
Last updated Sep 12, 2026
Free listing since Aug 26, 2026
Input text, image, video, pdf, reasoning
Output text
Capabilities reasoning, tool calling
Open weights Yes
AI Recommendation

Should you use glm-5.3-flash?

glm-5.3-flash is listed for chat workloads and supports a 1.3M context window.

Use it when NVIDIA NIM's free tier is enough for evaluation, demos, or light production traffic.

Best for
  • Chat
Strengths & Weaknesses

Strengths

  • Long context window
  • Reasoning mode listed
  • Tool calling support
  • Works with OpenAI-style SDKs

Watch outs

  • Free-tier rate limits apply
Benchmark Overview

Benchmark signals for glm-5.3-flash

Measured data
Intelligence General reasoning and instruction following
41.9/100
Coding Programming and code generation
71.5/100
Agentic Tool use and multi-step tasks
51.2/100
Speed Observed generation speed
95 tok/s
Context Maximum listed context window
1.3M
Pricing

glm-5.3-flash pricing per 1M tokens

Free tier listed
Input $0.15 per 1M tokens
Output $0.5 per 1M tokens
Access Available NVIDIA NIM
Rate limit Up to 40 RPM provider policy
Availability

glm-5.3-flash availability by provider

Current provider only

We only found this glm listing on NVIDIA NIM in the current catalog. Use the related models section for nearby alternatives.

Provider Model listing Access Context API Limits
NVIDIA NIM z-ai/glm-5.3-flash Free tier 1.3M OpenAI-style Up to 40 RPM
View NVIDIA NIM setup guide →
Typical Use Cases

glm-5.3-flash use cases

Chat

glm-5.3-flash is tagged for chat in this catalog and works with OpenAI-compatible client libraries.

FAQ

glm-5.3-flash free API FAQ

Is glm-5.3-flash free to use?

glm-5.3-flash is listed with free API access on NVIDIA NIM, subject to the provider's quota and account policy.

What is the glm-5.3-flash model ID?

The model ID shown in this catalog is z-ai/glm-5.3-flash.

What are the glm-5.3-flash free tier rate limits on NVIDIA NIM?

The listed free tier limit is Up to 40 RPM. Limits can change per account tier, so confirm against the provider dashboard.

What context window does glm-5.3-flash support?

The listed context window is 1.3M tokens with up to 131K output tokens.

More about glm-5.3-flash

z-ai/glm-5.3-flash — free model from NVIDIA NIM (z-ai).

For API keys, setup steps, and provider-level limits, see the NVIDIA NIM provider page.