NVIDIA NIM logo

Llama-3.3-70B-Instruct Free API on NVIDIA NIM

Free API Verified
★★★★★★★★★★ 3.5 Benchmark-backed score

Popular open Llama workhorse for multilingual chat, coding, and self-hosting

Free APIOpenAI compatibleTool callingJSON modeText
API Details

Copy-ready connection details

Use these values in your client or SDK.
Base URL
https://integrate.api.nvidia.com/v1
Model ID
meta/llama-3.3-70b-instruct
API format OpenAI Chat Completions
Technical Details

Llama-3.3-70B-Instruct specifications

Provider and model catalog
Context window 128K
Max output 4K
Status Online
Family llama
Knowledge cutoff 2023-12
Released Dec 6, 2024
Last updated Jul 7, 2026
Free listing since Dec 6, 2024
Input text
Output text
Capabilities tool calling, structured output, file attachments, temperature control
Open weights Yes

External benchmark references

Benchmark Score Metric Date
Artificial Analysis Coding Index 10.7 index 2026-03-11
SciCode 26 percent correct 2026-03-11
Terminal-Bench Hard 3 success rate 2026-03-11
AI Recommendation

Should you use Llama-3.3-70B-Instruct?

Llama-3.3-70B-Instruct is listed for chat workloads and supports a 128K context window.

Use it when NVIDIA NIM's free tier is enough for evaluation, demos, or light production traffic.

Best for
  • Chat
Strengths & Weaknesses

Strengths

  • Long context window
  • Tool calling support
  • Structured JSON output
  • Open weights available

Watch outs

  • Vision support is not listed
Benchmark Overview

Benchmark signals for Llama-3.3-70B-Instruct

Measured data
Intelligence General reasoning and instruction following
9.4/100
Coding Programming and code generation
11.9/100
Agentic Tool use and multi-step tasks
0.3/100
Context Maximum listed context window
128K
Pricing

Llama-3.3-70B-Instruct pricing per 1M tokens

Free tier listed
Input $0 per 1M tokens
Output $0 per 1M tokens
Free access Available NVIDIA NIM
Availability

Llama-3.3-70B-Instruct availability by provider

3 alternatives

We found 4 provider listings for llama-3-3-70b-instruct. Check model ID, quota, pricing, and API format before switching providers.

Provider Model listing Access Context API Limits
NVIDIA NIM Llama-3.3-70B-Instruct Free tier 128K Native Varies
Cloudflare Workers AI @cf/meta/llama-3.3-70b-instruct-fp8-fast Free tier 131K Native 10K neurons/day (shared)
GitHub Models Llama-3.3-70B-Instruct Free tier 131K OpenAI-style 15 RPM, 150 RPD
OVHcloud AI Endpoints Meta-Llama-3_3-70B-Instruct Free tier 131K OpenAI-style 2 RPM (anonymous)
View NVIDIA NIM setup guide →
Typical Use Cases

Llama-3.3-70B-Instruct use cases

Chat

Llama-3.3-70B-Instruct is tagged for chat in this catalog and works with OpenAI-compatible client libraries.

FAQ

Llama-3.3-70B-Instruct free API FAQ

Is Llama-3.3-70B-Instruct free to use?

Llama-3.3-70B-Instruct is listed with free API access on NVIDIA NIM, subject to the provider's quota and account policy.

What is the Llama-3.3-70B-Instruct model ID?

The model ID shown in this catalog is meta/llama-3.3-70b-instruct.

What context window does Llama-3.3-70B-Instruct support?

The listed context window is 128K tokens with up to 4K output tokens.

More about Llama-3.3-70B-Instruct

Llama 3.3 70B Instruct on OVHcloud AI Endpoints provides free access to Meta's flagship 70B model with 131K context and OpenAI-compatible API. OVHcloud is a major European cloud provider, making this endpoint particularly attractive for EU-based developers who want GDPR-compliant infrastructure. The free tier output is capped at 4K tokens per request — adequate for chat and short generation, but not long-form writing. Registration required.

For API keys, setup steps, and provider-level limits, see the NVIDIA NIM provider page.