Should you use Nemotron 3 Super 120B A12B?
Nemotron 3 Super 120B A12B is listed for chat workloads and supports a 262K context window.
Use it when NVIDIA NIM's free tier is enough for evaluation, demos, or light production traffic.
https://integrate.api.nvidia.com/v1 nvidia/nemotron-3-super-120b-a12b Nemotron 3 Super 120B A12B is listed for chat workloads and supports a 262K context window.
Use it when NVIDIA NIM's free tier is enough for evaluation, demos, or light production traffic.
We found 3 provider listings for nemotron-3-super-120b-a12b. Check model ID, quota, pricing, and API format before switching providers.
| Provider | Model listing | Access | Context | API | Limits |
|---|---|---|---|---|---|
| | Nemotron 3 Super 120B A12B | Free tier | 262K | Native | Varies |
| | NVIDIA: Nemotron 3 Super (free) | Free tier | 262K | OpenAI-style | 200 req/day (free tier) |
| | nvidia/nemotron-3-super-120b-a12b:free | Free tier | 262K | OpenAI-style | ~200 req/hr |
| Model | Provider | Context | Access |
|---|---|---|---|
| NVIDIA: Nemotron 3 Super (free) | OpenRouter | 262K | Free tier |
| nvidia/nemotron-3-super-120b-a12b:free | Kilo Code | 262K | Free tier |
| 01-ai/yi-large | NVIDIA NIM | 131K | Free tier |
| adept/fuyu-8b | NVIDIA NIM | 131K | Free tier |
| ai21labs/jamba-1.5-large-instruct | NVIDIA NIM | 131K | Free tier |
Nemotron 3 Super 120B A12B is tagged for chat in this catalog and works with OpenAI-compatible client libraries.
Nemotron 3 Super 120B A12B is listed with free API access on NVIDIA NIM, subject to the provider's quota and account policy.
The model ID shown in this catalog is nvidia/nemotron-3-super-120b-a12b.
The listed context window is 262K tokens with up to 262K output tokens.
NVIDIA Nemotron 3 Super 120B A12B is an open hybrid MoE model with 120B total parameters (12B active), available free on OpenRouter. Built on a hybrid Mamba-Transformer Mixture-of-Experts architecture with multi-token prediction (MTP), it delivers over 50% higher token generation speed compared to leading open models. Designed for multi-agent applications — long-term agent coherence, cross-document reasoning, and multi-step task planning. Trained with multi-environment RL across 10+ environments. Latent MoE calls 4 experts for the cost of one. Up to 1M context window. Fully open: weights, datasets, and recipes. OpenAI-compatible. Free tier: 200 RPD.
For API keys, setup steps, and provider-level limits, see the NVIDIA NIM provider page.