NVIDIA NIM logo

llama-nemotron-embed-vl-1b-v2 Free API on NVIDIA NIM

Free API Verified
Catalog metadata matched Reasoning models.dev metadata

Embedding model for semantic search, retrieval, clustering, and ranking pipelines

Free APIOpenAI compatibleReasoningEmbeddingTextImage
API Details

Copy-ready connection details

Use these values in your client or SDK.
Base URL
https://integrate.api.nvidia.com/v1
Model ID
nvidia/llama-nemotron-embed-vl-1b-v2
API format OpenAI Chat Completions
Technical Details

llama-nemotron-embed-vl-1b-v2 specifications

Provider and model catalog
Context window 131K
Max output 8K
Status Online
Family nemotron
Released Feb 10, 2026
Last updated Jul 5, 2026
Free listing since Feb 10, 2026
Input text, image
Output text
Capabilities reasoning, file attachments
Open weights Yes
AI Recommendation

Should you use llama-nemotron-embed-vl-1b-v2?

llama-nemotron-embed-vl-1b-v2 is listed for chat, reasoning, embedding workloads and supports a 131K context window.

Use it when NVIDIA NIM's free tier is enough for evaluation, demos, or light production traffic.

Best for
  • Chat
  • Reasoning
  • Embedding
Strengths & Weaknesses

Strengths

  • Strong reasoning profile
  • Long context window
  • Reasoning mode listed
  • Open weights available

Watch outs

  • Free-tier rate limits apply
  • Tool calling is not confirmed
Pricing

llama-nemotron-embed-vl-1b-v2 pricing per 1M tokens

Free tier listed
Input $0 per 1M tokens
Output $0 per 1M tokens
Free access Available NVIDIA NIM
Rate limit Up to 40 RPM provider policy
Typical Use Cases

llama-nemotron-embed-vl-1b-v2 use cases

Chat

llama-nemotron-embed-vl-1b-v2 is tagged for chat in this catalog and works with OpenAI-compatible client libraries.

Reasoning

llama-nemotron-embed-vl-1b-v2 is tagged for reasoning in this catalog and works with OpenAI-compatible client libraries.

Embedding

llama-nemotron-embed-vl-1b-v2 is tagged for embedding in this catalog and works with OpenAI-compatible client libraries.

FAQ

llama-nemotron-embed-vl-1b-v2 free API FAQ

Is llama-nemotron-embed-vl-1b-v2 free to use?

llama-nemotron-embed-vl-1b-v2 is listed with free API access on NVIDIA NIM, subject to the provider's quota and account policy.

What is the llama-nemotron-embed-vl-1b-v2 model ID?

The model ID shown in this catalog is nvidia/llama-nemotron-embed-vl-1b-v2.

What are the llama-nemotron-embed-vl-1b-v2 free tier rate limits on NVIDIA NIM?

The listed free tier limit is Up to 40 RPM. Limits can change per account tier, so confirm against the provider dashboard.

What context window does llama-nemotron-embed-vl-1b-v2 support?

The listed context window is 131K tokens with up to 8K output tokens.

More about llama-nemotron-embed-vl-1b-v2

Embedding model for semantic search, retrieval, clustering, and ranking pipelines

For API keys, setup steps, and provider-level limits, see the NVIDIA NIM provider page.