Gemini 2.5 Flash-Lite API Pricing: Google's Cheapest Model at $0.10/M Tokens
Gemini 2.5 Flash-Lite is Google's lowest-cost general-purpose model — 1M context, vision support, and batch API at 50% off.
TL;DR
- Price: $0.10/M input, $0.40/M output — Google's cheapest general-purpose model
- Batch API: $0.05/M input, $0.20/M output — 50% cheaper for async workloads
- Context: 1M tokens — same as Google's most expensive models, 7.8x larger than GPT-5 nano
- Vision: Yes — text, image, and video input supported. One of the cheapest multimodal models
- Provider: Google AI (ai.google.dev)
- Best for: High-volume classification, document processing, image analysis, translation, content moderation
- Trade-off: Less capable reasoning than Gemini 2.5 Flash ($0.30/$2.50) or Gemini 3.1 Pro ($2/$12)
Gemini 2.5 Flash-Lite Pricing Breakdown
At $0.10 per million input tokens and $0.40 per million output tokens, Gemini 2.5 Flash-Lite is Google's cheapest model. With batch API, the price drops to $0.05/$0.20 — making it one of the cheapest ways to process large volumes of text or images.
| Monthly Volume | Standard Input | Standard Output | Batch Input | Batch Output |
|---|---|---|---|---|
| 1M tokens | $0.10 | $0.40 | $0.05 | $0.20 |
| 10M tokens | $1.00 | $4.00 | $0.50 | $2.00 |
| 100M tokens | $10.00 | $40.00 | $5.00 | $20.00 |
| 1B tokens | $100.00 | $400.00 | $50.00 | $200.00 |
At 100M tokens/month with batch API, Gemini 2.5 Flash-Lite costs just $25. The same volume on GPT-5.4 nano would cost $350, and on Claude Haiku 4.5 it would cost $1,500.
How Gemini 2.5 Flash-Lite Compares to Other Budget Models
| Model | Input $/M | Output $/M | Context | Vision | Provider |
|---|---|---|---|---|---|
| Qwen 3.7 Flash | $0.03 | $0.13 | 1M | ✅ | Alibaba |
| GPT-5 nano | $0.05 | $0.40 | 128K | ❌ | OpenAI |
| GPT-oss 20B | $0.08 | $0.35 | 128K | ❌ | OpenAI (self-host) |
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | 1M | ✅ | |
| Ministral 3 3B | $0.10 | $0.10 | 128K | ❌ | Mistral |
| DeepSeek V4 Flash | $0.14 | $0.28 | 1M | ❌ | DeepSeek |
| Mistral Small 4 | $0.15 | $0.60 | 128K | ❌ | Mistral |
Gemini 2.5 Flash-Lite is the cheapest model with both 1M context AND vision support from a major US provider. Only Qwen 3.7 Flash ($0.03/$0.13) is cheaper, but that's from Alibaba — some enterprises prefer Google's infrastructure and data residency options.
Google Flash-Lite Family: Which Tier to Choose
Google offers three Flash-Lite tiers. Here's how they compare:
| Model | Input $/M | Output $/M | Context | Best For |
|---|---|---|---|---|
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | 1M | Cheapest option, batch workloads |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 | 1M | Better quality, still budget |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | 1M | Latest generation, best quality |
Rule of thumb: Use Gemini 2.5 Flash-Lite for high-volume, cost-sensitive tasks. Upgrade to 3.1 Flash-Lite when you need better reasoning or more accurate extraction. Use 3.5 Flash-Lite for tasks where quality matters most but you still want budget pricing.
Batch API: 50% Off for Async Workloads
Gemini 2.5 Flash-Lite supports batch API at half price — $0.05/M input and $0.20/M output. This is ideal for:
- Overnight processing: Run large classification or extraction jobs while you sleep
- Data pipeline backfills: Reprocess historical data at 50% cost
- Report generation: Batch summarize documents, emails, or tickets
- Training data preparation: Label, classify, or extract features at scale
At batch pricing, Gemini 2.5 Flash-Lite is cheaper than Qwen 3.7 Flash ($0.03/$0.13) on output tokens — $0.20 vs $0.13 is close, but you get 1M context and Google's infrastructure.
When to Use Gemini 2.5 Flash-Lite
📊 Classification
Sentiment analysis, intent detection, content categorization. Batch API makes high-volume classification extremely cheap.
📄 Document Processing
Parse invoices, contracts, or reports up to 1M tokens. Extract structured data from long documents in a single call.
🖼️ Image Analysis
Describe images, extract text from screenshots, classify visual content. One of the cheapest vision models available.
🛡️ Content Moderation
Flag inappropriate text and images. Process millions of items at minimal cost with batch API.
🌐 Translation
Translate documents, product listings, or user-generated content. Good quality for the price, especially with batch API.
📝 Summarization
Summarize long documents, meeting transcripts, or support tickets. 1M context handles even very long inputs.
When NOT to Use Gemini 2.5 Flash-Lite
Gemini 2.5 Flash-Lite is optimized for speed and cost, not complexity. Upgrade when you need:
- Complex reasoning: Multi-step logic, math, or code generation → use Gemini 2.5 Flash ($0.30/$2.50) or Claude Sonnet 5 ($2/$10)
- Nuanced writing: Creative content, marketing copy, or tone-sensitive communication → use Claude Opus 5 ($5/$25)
- Agentic workflows: Tool use, multi-step planning, or autonomous decision-making → use Gemini 3.1 Pro ($2/$12) or Claude Sonnet 5
- Cheapest possible: If you need the absolute lowest price → use Qwen 3.7 Flash ($0.03/$0.13)
Real-World Cost Scenario
Let's say you're building a document processing pipeline that processes 50,000 documents per month, with an average of 2,000 input tokens and 500 output tokens per document:
| Model | Monthly Input | Monthly Output | Total |
|---|---|---|---|
| Qwen 3.7 Flash | $3.00 | $3.25 | $6.25 |
| Gemini 2.5 Flash-Lite (Batch) | $5.00 | $5.00 | $10.00 |
| Gemini 2.5 Flash-Lite (Standard) | $10.00 | $10.00 | $20.00 |
| GPT-5 nano | $5.00 | $10.00 | $15.00 |
| Gemini 3.1 Flash-Lite | $25.00 | $37.50 | $62.50 |
| Claude Haiku 4.5 | $100.00 | $125.00 | $225.00 |
At batch pricing, Gemini 2.5 Flash-Lite costs $10/month for 50K documents — just 60% more than Qwen 3.7 Flash, but with Google's infrastructure and 1M context window.
Frequently Asked Questions
Calculate Your Gemini 2.5 Flash-Lite Costs
Enter your token counts and see exactly what you'd pay. Compare against 93 other models instantly.