NVIDIA NIM logo

llama-3.2-nemoretriever-1b-vlm-embed-v1 Free API on NVIDIA NIM

Free API Verified
Catalog profile API ready provider catalog metadata

nvidia/llama-3.2-nemoretriever-1b-vlm-embed-v1 — free model from NVIDIA NIM (nvidia).

Free APIOpenAI compatibleEmbeddingRerank
API Details

Copy-ready connection details

Use these values in your client or SDK.
Base URL
https://integrate.api.nvidia.com/v1
Model ID
nvidia/llama-3.2-nemoretriever-1b-vlm-embed-v1
API format OpenAI Chat Completions
Technical Details

llama-3.2-nemoretriever-1b-vlm-embed-v1 specifications

Provider catalog
Context window 131K
Max output 8K
Status Online
Last updated Jun 30, 2026
Free listing since Jun 17, 2026
Input embedding, rerank
Output text
AI Recommendation

Should you use llama-3.2-nemoretriever-1b-vlm-embed-v1?

llama-3.2-nemoretriever-1b-vlm-embed-v1 is listed for chat, embedding workloads and supports a 131K context window.

Use it when NVIDIA NIM's free tier is enough for evaluation, demos, or light production traffic.

Best for
  • Chat
  • Embedding
Strengths & Weaknesses

Strengths

  • Long context window
  • Works with OpenAI-style SDKs
  • Live API verification available

Watch outs

  • Free-tier rate limits apply
  • Vision support is not listed
  • Tool calling is not confirmed
Typical Use Cases

llama-3.2-nemoretriever-1b-vlm-embed-v1 use cases

Chat

llama-3.2-nemoretriever-1b-vlm-embed-v1 is tagged for chat in this catalog and works with OpenAI-compatible client libraries.

Embedding

llama-3.2-nemoretriever-1b-vlm-embed-v1 is tagged for embedding in this catalog and works with OpenAI-compatible client libraries.

FAQ

llama-3.2-nemoretriever-1b-vlm-embed-v1 free API FAQ

Is llama-3.2-nemoretriever-1b-vlm-embed-v1 free to use?

llama-3.2-nemoretriever-1b-vlm-embed-v1 is listed with free API access on NVIDIA NIM, subject to the provider's quota and account policy.

What is the llama-3.2-nemoretriever-1b-vlm-embed-v1 model ID?

The model ID shown in this catalog is nvidia/llama-3.2-nemoretriever-1b-vlm-embed-v1.

What are the llama-3.2-nemoretriever-1b-vlm-embed-v1 free tier rate limits on NVIDIA NIM?

The listed free tier limit is Up to 40 RPM. Limits can change per account tier, so confirm against the provider dashboard.

What context window does llama-3.2-nemoretriever-1b-vlm-embed-v1 support?

The listed context window is 131K tokens with up to 8K output tokens.

More about llama-3.2-nemoretriever-1b-vlm-embed-v1

nvidia/llama-3.2-nemoretriever-1b-vlm-embed-v1 — free model from NVIDIA NIM (nvidia).

For API keys, setup steps, and provider-level limits, see the NVIDIA NIM provider page.