Should you use gemma-4-26b-a4b-it?
gemma-4-26b-a4b-it is listed for chat workloads and supports a 256K context window.
Use it when Cloudflare Workers AI's free tier is enough for evaluation, demos, or light production traffic.
https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/run @cf/google/gemma-4-26b-a4b-it gemma-4-26b-a4b-it is listed for chat workloads and supports a 256K context window.
Use it when Cloudflare Workers AI's free tier is enough for evaluation, demos, or light production traffic.
We found 3 provider listings for gemma-4-26b-a4b-it. Check model ID, quota, pricing, and API format before switching providers.
| Provider | Model listing | Access | Context | API | Limits |
|---|---|---|---|---|---|
| | @cf/google/gemma-4-26b-a4b-it | Free tier | 256K | Native | 10K neurons/day (shared) |
| | Google: Gemma 4 26B A4B (free) | Free tier | 262K | OpenAI-style | 200 req/day (free tier) |
| | Gemma 4 26B A4B IT | Free tier | 262K | Native | Varies |
| Model | Provider | Context | Access |
|---|---|---|---|
| Google: Gemma 4 26B A4B (free) | OpenRouter | 262K | Free tier |
| Gemma 4 26B A4B IT | Google Gemini | 262K | Free tier |
| Mistral 7B | Cloudflare Workers AI | 33K | Free tier |
| Qwen 1.5 7B | Cloudflare Workers AI | 33K | Free tier |
| @cf/meta/llama-3.3-70b-instruct-fp8-fast | Cloudflare Workers AI | 131K | Free tier |
gemma-4-26b-a4b-it is tagged for chat in this catalog and works with OpenAI-compatible client libraries.
gemma-4-26b-a4b-it is listed with free API access on Cloudflare Workers AI, subject to the provider's quota and account policy.
The model ID shown in this catalog is @cf/google/gemma-4-26b-a4b-it.
The listed free tier limit is 10K neurons/day (shared). Limits can change per account tier, so confirm against the provider dashboard.
The listed context window is 256K tokens with up to 131K output tokens.
Gemma 4 26B is Google's latest open model, running on Cloudflare Workers AI with a 256K context window and a 26B-parameter architecture using 4 active experts (MoE). It brings Google's latest advances in small-model efficiency to the edge, making it suitable for long-form text tasks, summarization, and instruction-following workloads. The 256K context window is particularly useful for processing full-length documents or long conversation histories. Operates on Cloudflare's shared 10,000 Neurons/day free allocation; API access uses Cloudflare's native format, not the OpenAI SDK.
For API keys, setup steps, and provider-level limits, see the Cloudflare Workers AI provider page.