Gemini 2.5 Flash-Lite API Pricing: Google's Cheapest Model at $0.10/M Tokens

Gemini 2.5 Flash-Lite is Google's lowest-cost general-purpose model — 1M context, vision support, and batch API at 50% off.

Updated Aug 9, 2026 · 93 models tracked across 11 providers

TL;DR

Gemini 2.5 Flash-Lite Pricing Breakdown

At $0.10 per million input tokens and $0.40 per million output tokens, Gemini 2.5 Flash-Lite is Google's cheapest model. With batch API, the price drops to $0.05/$0.20 — making it one of the cheapest ways to process large volumes of text or images.

Monthly Volume Standard Input Standard Output Batch Input Batch Output
1M tokens$0.10$0.40$0.05$0.20
10M tokens$1.00$4.00$0.50$2.00
100M tokens$10.00$40.00$5.00$20.00
1B tokens$100.00$400.00$50.00$200.00

At 100M tokens/month with batch API, Gemini 2.5 Flash-Lite costs just $25. The same volume on GPT-5.4 nano would cost $350, and on Claude Haiku 4.5 it would cost $1,500.

How Gemini 2.5 Flash-Lite Compares to Other Budget Models

Model Input $/M Output $/M Context Vision Provider
Qwen 3.7 Flash $0.03 $0.13 1M Alibaba
GPT-5 nano $0.05 $0.40 128K OpenAI
GPT-oss 20B $0.08 $0.35 128K OpenAI (self-host)
Gemini 2.5 Flash-Lite $0.10 $0.40 1M Google
Ministral 3 3B $0.10 $0.10 128K Mistral
DeepSeek V4 Flash $0.14 $0.28 1M DeepSeek
Mistral Small 4 $0.15 $0.60 128K Mistral

Gemini 2.5 Flash-Lite is the cheapest model with both 1M context AND vision support from a major US provider. Only Qwen 3.7 Flash ($0.03/$0.13) is cheaper, but that's from Alibaba — some enterprises prefer Google's infrastructure and data residency options.

Google Flash-Lite Family: Which Tier to Choose

Google offers three Flash-Lite tiers. Here's how they compare:

ModelInput $/MOutput $/MContextBest For
Gemini 2.5 Flash-Lite$0.10$0.401MCheapest option, batch workloads
Gemini 3.1 Flash-Lite$0.25$1.501MBetter quality, still budget
Gemini 3.5 Flash-Lite$0.30$2.501MLatest generation, best quality

Rule of thumb: Use Gemini 2.5 Flash-Lite for high-volume, cost-sensitive tasks. Upgrade to 3.1 Flash-Lite when you need better reasoning or more accurate extraction. Use 3.5 Flash-Lite for tasks where quality matters most but you still want budget pricing.

Batch API: 50% Off for Async Workloads

Gemini 2.5 Flash-Lite supports batch API at half price — $0.05/M input and $0.20/M output. This is ideal for:

At batch pricing, Gemini 2.5 Flash-Lite is cheaper than Qwen 3.7 Flash ($0.03/$0.13) on output tokens — $0.20 vs $0.13 is close, but you get 1M context and Google's infrastructure.

When to Use Gemini 2.5 Flash-Lite

📊 Classification

Sentiment analysis, intent detection, content categorization. Batch API makes high-volume classification extremely cheap.

📄 Document Processing

Parse invoices, contracts, or reports up to 1M tokens. Extract structured data from long documents in a single call.

🖼️ Image Analysis

Describe images, extract text from screenshots, classify visual content. One of the cheapest vision models available.

🛡️ Content Moderation

Flag inappropriate text and images. Process millions of items at minimal cost with batch API.

🌐 Translation

Translate documents, product listings, or user-generated content. Good quality for the price, especially with batch API.

📝 Summarization

Summarize long documents, meeting transcripts, or support tickets. 1M context handles even very long inputs.

When NOT to Use Gemini 2.5 Flash-Lite

Gemini 2.5 Flash-Lite is optimized for speed and cost, not complexity. Upgrade when you need:

Real-World Cost Scenario

Let's say you're building a document processing pipeline that processes 50,000 documents per month, with an average of 2,000 input tokens and 500 output tokens per document:

ModelMonthly InputMonthly OutputTotal
Qwen 3.7 Flash$3.00$3.25$6.25
Gemini 2.5 Flash-Lite (Batch)$5.00$5.00$10.00
Gemini 2.5 Flash-Lite (Standard)$10.00$10.00$20.00
GPT-5 nano$5.00$10.00$15.00
Gemini 3.1 Flash-Lite$25.00$37.50$62.50
Claude Haiku 4.5$100.00$125.00$225.00

At batch pricing, Gemini 2.5 Flash-Lite costs $10/month for 50K documents — just 60% more than Qwen 3.7 Flash, but with Google's infrastructure and 1M context window.

Frequently Asked Questions

How much does Gemini 2.5 Flash-Lite cost?
$0.10 per million input tokens and $0.40 per million output tokens. Batch API pricing is 50% cheaper at $0.05/$0.20. This makes it Google's cheapest general-purpose model.
What is the context window of Gemini 2.5 Flash-Lite?
1M tokens — the same as Google's most expensive models. This is 7.8x larger than GPT-5 nano's 128K window, at only 2x the price.
Does Gemini 2.5 Flash-Lite support vision?
Yes. Gemini 2.5 Flash-Lite supports text, image, and video input. It's one of the cheapest multimodal models available — only Qwen 3.7 Flash ($0.03/$0.13) is cheaper for vision tasks.
Is Gemini 2.5 Flash-Lite cheaper than GPT-5 nano?
GPT-5 nano ($0.05/$0.40) is cheaper on input (2x) but the same on output. However, Gemini 2.5 Flash-Lite offers 1M context (vs 128K), vision support, and batch API at $0.05/$0.20 — making it cheaper at scale for multimodal or long-context workloads.
Is Gemini 2.5 Flash-Lite being deprecated?
No. Gemini 2.5 Flash-Lite is a current stable model. Google has deprecated Gemini 2.0 Flash and 2.0 Flash Lite (replaced by 3.x versions), but 2.5 Flash-Lite remains available.
What can Gemini 2.5 Flash-Lite be used for?
Classification, data extraction, document processing, image analysis, content moderation, translation, and summarization — any high-volume task where you need 1M context or vision at the lowest Google price. For complex reasoning, upgrade to Gemini 2.5 Flash or Gemini 3.1 Pro.

Calculate Your Gemini 2.5 Flash-Lite Costs

Enter your token counts and see exactly what you'd pay. Compare against 93 other models instantly.