NVIDIA NIM logo

llama-nemotron-embed-1b-v2 Free API on NVIDIA NIM

Free API Verified
Catalog profile Reasoning provider catalog metadata

nvidia/llama-nemotron-embed-1b-v2 — free model from NVIDIA NIM (nvidia).

Free APIOpenAI compatibleReasoningEmbeddingTextImage
API Details

Copy-ready connection details

Use these values in your client or SDK.
Base URL
https://integrate.api.nvidia.com/v1
Model ID
nvidia/llama-nemotron-embed-1b-v2
API format OpenAI Chat Completions
Technical Details

llama-nemotron-embed-1b-v2 specifications

Provider catalog
Context window 131K
Max output 8K
Status Online
Family nemotron
Released Feb 10, 2026
Last updated Jun 30, 2026
Free listing since Feb 10, 2026
Input embedding, text, image
Output text
Capabilities reasoning
Open weights Yes
AI Recommendation

Should you use llama-nemotron-embed-1b-v2?

llama-nemotron-embed-1b-v2 is listed for chat, reasoning, embedding workloads and supports a 131K context window.

Use it when NVIDIA NIM's free tier is enough for evaluation, demos, or light production traffic.

Best for
  • Chat
  • Reasoning
  • Embedding
Strengths & Weaknesses

Strengths

  • Strong reasoning profile
  • Long context window
  • Reasoning mode listed
  • Works with OpenAI-style SDKs

Watch outs

  • Free-tier rate limits apply
  • Tool calling is not confirmed
Availability

llama-nemotron-embed-1b-v2 availability by provider

Current provider only

We only found this nemotron listing on NVIDIA NIM in the current catalog. Use the related models section for nearby alternatives.

Provider Model listing Access Context API Limits
NVIDIA NIM nvidia/llama-nemotron-embed-1b-v2 Free tier 131K OpenAI-style Up to 40 RPM
View NVIDIA NIM setup guide →
Typical Use Cases

llama-nemotron-embed-1b-v2 use cases

Chat

llama-nemotron-embed-1b-v2 is tagged for chat in this catalog and works with OpenAI-compatible client libraries.

Reasoning

llama-nemotron-embed-1b-v2 is tagged for reasoning in this catalog and works with OpenAI-compatible client libraries.

Embedding

llama-nemotron-embed-1b-v2 is tagged for embedding in this catalog and works with OpenAI-compatible client libraries.

FAQ

llama-nemotron-embed-1b-v2 free API FAQ

Is llama-nemotron-embed-1b-v2 free to use?

llama-nemotron-embed-1b-v2 is listed with free API access on NVIDIA NIM, subject to the provider's quota and account policy.

What is the llama-nemotron-embed-1b-v2 model ID?

The model ID shown in this catalog is nvidia/llama-nemotron-embed-1b-v2.

What are the llama-nemotron-embed-1b-v2 free tier rate limits on NVIDIA NIM?

The listed free tier limit is Up to 40 RPM. Limits can change per account tier, so confirm against the provider dashboard.

What context window does llama-nemotron-embed-1b-v2 support?

The listed context window is 131K tokens with up to 8K output tokens.

More about llama-nemotron-embed-1b-v2

nvidia/llama-nemotron-embed-1b-v2 — free model from NVIDIA NIM (nvidia).

For API keys, setup steps, and provider-level limits, see the NVIDIA NIM provider page.