LLM7.io logo

gemini-2.5-flash-lite API status on LLM7.io

Check provider
★★★★★★★★★★ 3.5 Benchmark-backed score

Lean Gemini 2.5 lane for cheap multimodal traffic and quick agents

Check providerOpenAI compatibleReasoningTool callingJSON modePDF inputText
API Details

Copy-ready connection details

Use these values in your client or SDK.
Base URL
https://api.llm7.io/v1
Model ID
gemini-2-5-flash-lite
API format OpenAI Chat Completions
Technical Details

gemini-2.5-flash-lite specifications

Provider and model catalog
Context window 131K
Max output 131K
Status Check provider
Family gemini-flash-lite
Knowledge cutoff 2025-01
Released Jun 17, 2025
Last updated Jul 30, 2026
Free listing since Jun 17, 2025
Input text, image, audio, video, pdf
Output text
Capabilities reasoning, tool calling, structured output, file attachments, temperature control

External benchmark references

Benchmark Score Metric Date
Artificial Analysis Coding Index 9.5 index 2026-03-11
SciCode 19.3 percent correct 2026-03-11
Terminal-Bench Hard 4.5 success rate 2026-03-11
AI Recommendation

Should you use gemini-2.5-flash-lite?

gemini-2.5-flash-lite is listed for chat workloads and supports a 131K context window.

Check LLM7.io's current endpoint status before integrating this listing. Use the alternatives table if the provider listing is unavailable.

Best for
  • Chat
Strengths & Weaknesses

Strengths

  • Long context window
  • Reasoning mode listed
  • Tool calling support
  • Structured JSON output

Watch outs

  • Current endpoint availability needs provider confirmation
  • Free-tier rate limits apply
Benchmark Overview

Benchmark signals for gemini-2.5-flash-lite

Measured data
Intelligence General reasoning and instruction following
17.6/100
Coding Programming and code generation
18.2/100
Agentic Tool use and multi-step tasks
11.7/100
Speed Observed generation speed
215 tok/s
Context Maximum listed context window
131K
Pricing

gemini-2.5-flash-lite pricing per 1M tokens

Check provider
Input $0.1 per 1M tokens
Output $0.4 per 1M tokens
Free access Check provider LLM7.io
Rate limit 30 RPM (120 with token) provider policy
Typical Use Cases

gemini-2.5-flash-lite use cases

Chat

gemini-2.5-flash-lite is tagged for chat in this catalog and works with OpenAI-compatible client libraries.

FAQ

gemini-2.5-flash-lite free API FAQ

Is gemini-2.5-flash-lite free to use?

gemini-2.5-flash-lite appears in the free model catalog for LLM7.io, but its current endpoint availability should be confirmed with the provider before use.

What is the gemini-2.5-flash-lite model ID?

The model ID shown in this catalog is gemini-2-5-flash-lite.

What are the gemini-2.5-flash-lite free tier rate limits on LLM7.io?

The listed free tier limit is 30 RPM (120 with token). Limits can change per account tier, so confirm against the provider dashboard.

What context window does gemini-2.5-flash-lite support?

The listed context window is 131K tokens with up to 131K output tokens.

More about gemini-2.5-flash-lite

The Google Gemini 2.5 Flash-Lite is a pre-trained 1M token text-based LLM ideal for chat applications, producing up to 65k tokens at 15 RPM and 1k RPD. It's open to developers but not OpenAI compatible or requiring a credit card.

For API keys, setup steps, and provider-level limits, see the LLM7.io provider page.