Should you use step-3.7-flash?
step-3.7-flash is listed for chat workloads and supports a 262K context window.
Use it when Kilo Code's free tier is enough for evaluation, demos, or light production traffic.
https://api.kilo.ai/api/gateway stepfun/step-3.7-flash:free | Benchmark | Score | Metric | Date |
|---|---|---|---|
| SWE-Bench Pro | 56.3 | resolve rate | 2026-05-29 |
| SWE-Bench Verified | 76.5 | resolved | 2026-05-29 |
| Terminal-Bench | 59.6 | success rate | 2026-05-29 |
| Humanity's Last Exam | 47.2 | accuracy | 2026-05-29 |
step-3.7-flash is listed for chat workloads and supports a 262K context window.
Use it when Kilo Code's free tier is enough for evaluation, demos, or light production traffic.
We only found this step-3.7-flash listing on Kilo Code in the current catalog. Use the related models section for nearby alternatives.
| Provider | Model listing | Access | Context | API | Limits |
|---|---|---|---|---|---|
| | stepfun/step-3.7-flash:free | Free tier | 262K | OpenAI-style | ~200 req/hr |
| Model | Provider | Context | Access |
|---|---|---|---|
| nvidia/nemotron-3-ultra-550b-a55b:free | Kilo Code | 1.0M | Free tier |
| nvidia/nemotron-3-super-120b-a12b:free | Kilo Code | 262K | Free tier |
| nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free | Kilo Code | 256K | Free tier |
| inclusionai/ling-3.0-flash:free | Kilo Code | 262K | Free tier |
| poolside/laguna-s-2.1:free | Kilo Code | 262K | Free tier |
step-3.7-flash is tagged for chat in this catalog and works with OpenAI-compatible client libraries.
step-3.7-flash is listed with free API access on Kilo Code, subject to the provider's quota and account policy.
The model ID shown in this catalog is stepfun/step-3.7-flash:free.
The listed free tier limit is ~200 req/hr. Limits can change per account tier, so confirm against the provider dashboard.
The listed context window is 262K tokens with up to 262K output tokens.
Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters per token. The model supports a 256K context window and exposes selectable reasoning levels (high/medium/low), letting callers trade off speed, cost, and depth of reasoning. Designed for coding, agentic workflows, structured outputs, and long-context productivity tasks.
For API keys, setup steps, and provider-level limits, see the Kilo Code provider page.