Should you use Inkling?
Inkling is listed for chat workloads and supports a 256K context window.
Use it when NVIDIA NIM's free tier is enough for evaluation, demos, or light production traffic.
Inkling is listed for chat workloads and supports a 256K context window.
Use it when NVIDIA NIM's free tier is enough for evaluation, demos, or light production traffic.
We only found this Inkling listing on NVIDIA NIM in the current catalog. Use the related models section for nearby alternatives.
| Provider | Model listing | Access | Context | API | Limits |
|---|---|---|---|---|---|
| | Inkling | Free tier | 256K | Native | Varies |
| Model | Provider | Context | Access |
|---|---|---|---|
| 01-ai/yi-large | NVIDIA NIM | 131K | Free tier |
| adept/fuyu-8b | NVIDIA NIM | 131K | Free tier |
| ai21labs/jamba-1.5-large-instruct | NVIDIA NIM | 131K | Free tier |
| aisingapore/sea-lion-7b-instruct | NVIDIA NIM | 131K | Free tier |
| baai/bge-m3 | NVIDIA NIM | 131K | Free tier |
Inkling is tagged for chat in this catalog and works with OpenAI-compatible client libraries.
Inkling is listed with free API access on NVIDIA NIM, subject to the provider's quota and account policy.
The model ID shown in this catalog is thinkingmachines/inkling.
The listed context window is 256K tokens with up to 256K output tokens.
Free Inkling API.
For API keys, setup steps, and provider-level limits, see the NVIDIA NIM provider page.