NVIDIA NIM logo

deepseek-v4-flash-0731 Free API on NVIDIA NIM

Free API Verified
★★★★★★★★★★ 3.5 Benchmark-backed score

Official DeepSeek V4 Flash release with enhanced agentic capabilities and integrated DSpark speculative decoding

Free APIOpenAI compatibleReasoningTool callingJSON modeTextReasoning
API Details

Copy-ready connection details

Use these values in your client or SDK.
Base URL
https://integrate.api.nvidia.com/v1
Model ID
deepseek-ai/deepseek-v4-flash-0731
API format OpenAI Chat Completions
Technical Details

deepseek-v4-flash-0731 specifications

Provider and model catalog
Context window 1.3M
Max output 944K
Status Online
Family deepseek-flash
Knowledge cutoff 2025-05
Released Jul 31, 2026
Last updated Sep 9, 2026
Free listing since Jul 31, 2026
Input text
Output text
Capabilities reasoning, tool calling, structured output, temperature control
Open weights Yes

External benchmark references

Benchmark Score Metric Date
Terminal-Bench 82.7 pass@1 Not listed
NL2Repo 54.2 resolve rate 2026-07-31
CyberGym 76.7 score 2026-07-31
DeepSWE 54.4 resolve rate 2026-07-31
AI Recommendation

Should you use deepseek-v4-flash-0731?

deepseek-v4-flash-0731 is listed for chat workloads and supports a 1.3M context window.

Use it when NVIDIA NIM's free tier is enough for evaluation, demos, or light production traffic.

Best for
  • Chat
Strengths & Weaknesses

Strengths

  • Long context window
  • Reasoning mode listed
  • Tool calling support
  • Structured JSON output

Watch outs

  • Free-tier rate limits apply
  • Vision support is not listed
Benchmark Overview

Benchmark signals for deepseek-v4-flash-0731

Measured data
Intelligence General reasoning and instruction following
34.5/100
Coding Programming and code generation
69.1/100
Agentic Tool use and multi-step tasks
41.7/100
Speed Observed generation speed
128 tok/s
Context Maximum listed context window
1.3M
Pricing

deepseek-v4-flash-0731 pricing per 1M tokens

Free tier listed
Input $0.44 per 1M tokens
Output $1.32 per 1M tokens
Free access Available NVIDIA NIM
Rate limit Up to 40 RPM provider policy
Availability

deepseek-v4-flash-0731 availability by provider

Current provider only

We only found this deepseek-v4-flash-0731 listing on NVIDIA NIM in the current catalog. Use the related models section for nearby alternatives.

Provider Model listing Access Context API Limits
NVIDIA NIM deepseek-ai/deepseek-v4-flash-0731 Free tier 1.3M OpenAI-style Up to 40 RPM
View NVIDIA NIM setup guide →
Typical Use Cases

deepseek-v4-flash-0731 use cases

Chat

deepseek-v4-flash-0731 is tagged for chat in this catalog and works with OpenAI-compatible client libraries.

FAQ

deepseek-v4-flash-0731 free API FAQ

Is deepseek-v4-flash-0731 free to use?

deepseek-v4-flash-0731 is listed with free API access on NVIDIA NIM, subject to the provider's quota and account policy.

What is the deepseek-v4-flash-0731 model ID?

The model ID shown in this catalog is deepseek-ai/deepseek-v4-flash-0731.

What are the deepseek-v4-flash-0731 free tier rate limits on NVIDIA NIM?

The listed free tier limit is Up to 40 RPM. Limits can change per account tier, so confirm against the provider dashboard.

What context window does deepseek-v4-flash-0731 support?

The listed context window is 1.3M tokens with up to 944K output tokens.

More about deepseek-v4-flash-0731

Official DeepSeek V4 Flash release with enhanced agentic capabilities and integrated DSpark speculative decoding

For API keys, setup steps, and provider-level limits, see the NVIDIA NIM provider page.