Should you use GPT OSS 20B?
GPT OSS 20B is listed for chat workloads and supports a 131K context window.
Use it when NVIDIA NIM's free tier is enough for evaluation, demos, or light production traffic.
https://integrate.api.nvidia.com/v1 openai/gpt-oss-20b GPT OSS 20B is listed for chat workloads and supports a 131K context window.
Use it when NVIDIA NIM's free tier is enough for evaluation, demos, or light production traffic.
We only found this GPT OSS 20B listing on NVIDIA NIM in the current catalog. Use the related models section for nearby alternatives.
| Provider | Model listing | Access | Context | API | Limits |
|---|---|---|---|---|---|
| | GPT OSS 20B | Free tier | 131K | Native | Varies |
| Model | Provider | Context | Access |
|---|---|---|---|
| 01-ai/yi-large | NVIDIA NIM | 131K | Free tier |
| adept/fuyu-8b | NVIDIA NIM | 131K | Free tier |
| ai21labs/jamba-1.5-large-instruct | NVIDIA NIM | 131K | Free tier |
| aisingapore/sea-lion-7b-instruct | NVIDIA NIM | 131K | Free tier |
| baai/bge-m3 | NVIDIA NIM | 131K | Free tier |
GPT OSS 20B is tagged for chat in this catalog and works with OpenAI-compatible client libraries.
GPT OSS 20B is listed with free API access on NVIDIA NIM, subject to the provider's quota and account policy.
The model ID shown in this catalog is openai/gpt-oss-20b.
The listed context window is 131K tokens with up to 33K output tokens.
GPT-OSS 20B is OpenAI's open-weight 21B-parameter Mixture-of-Experts model with 3.6B active per forward pass, available free on OpenRouter (also on Groq, Cerebras, and Cloudflare). Supports reasoning level configuration, fine-tuning, and agentic capabilities — function calling, tool use, and structured outputs. Trained in OpenAI's Harmony response format. Released under Apache 2.0. The low active parameter count enables lower-latency inference on consumer or single-GPU hardware. 131K context window. Text-only. OpenAI-compatible. Free tier: 200 RPD on OpenRouter.
For API keys, setup steps, and provider-level limits, see the NVIDIA NIM provider page.