Gemini Flash-Lite API Pricing: All 3 Google Budget Tiers Compared
Google offers three Flash-Lite models from $0.10/M to $0.30/M — all with 1M context and vision support. Here's which tier to pick.
TL;DR
- Cheapest: Gemini 2.5 Flash-Lite at $0.10/M input, $0.40/M output — best for high-volume simple tasks
- Mid-tier: Gemini 3.1 Flash-Lite at $0.25/M input, $1.50/M output — better reasoning, same 1M context
- Newest: Gemini 3.5 Flash-Lite at $0.30/M input, $2.50/M output — strongest Flash-Lite capabilities
- All three: 1M context window, vision support, Google AI API access
- Best for: Budget-conscious teams needing large context and vision without premium pricing
The 3 Gemini Flash-Lite Tiers
Gemini 2.5 Flash-Lite
$0.40/M output · 1M context
- Cheapest Flash-Lite
- Vision support
- Best for volume tasks
Gemini 3.1 Flash-Lite
$1.50/M output · 1M context
- Better reasoning
- Vision support
- Good balance
Gemini 3.5 Flash-Lite
$2.50/M output · 1M context
- Strongest capabilities
- Vision support
- Latest architecture
Flash-Lite Pricing Comparison
| Model | Input $/M | Output $/M | Context | Vision | Generation |
|---|---|---|---|---|---|
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | 1M | ✅ | 2.5 |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 | 1M | ✅ | 3.1 |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | 1M | ✅ | 3.5 |
The price gap between tiers is significant: 3.1 is 2.5x more expensive on input than 2.5, and 3.5 is 3x more. For most budget workloads, 2.5 Flash-Lite offers the best cost-per-token.
Monthly Cost at Different Volumes
| Monthly Volume | 2.5 Flash-Lite | 3.1 Flash-Lite | 3.5 Flash-Lite |
|---|---|---|---|
| 1M tokens | $0.50 | $1.75 | $2.80 |
| 10M tokens | $5.00 | $17.50 | $28.00 |
| 100M tokens | $50.00 | $175.00 | $280.00 |
| 1B tokens | $500.00 | $1,750.00 | $2,800.00 |
At 100M tokens/month, choosing 2.5 over 3.5 saves $230/month. The cost difference is negligible for small volumes but compounds fast at scale.
Flash-Lite vs Other Budget Models
| Model | Input $/M | Output $/M | Context | Vision | Provider |
|---|---|---|---|---|---|
| Qwen 3.7 Flash | $0.03 | $0.13 | 1M | ✅ | Alibaba |
| GPT-5 nano | $0.05 | $0.40 | 128K | ❌ | OpenAI |
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | 1M | ✅ | |
| DeepSeek V4 Flash | $0.14 | $0.28 | 1M | ❌ | DeepSeek |
| GPT-5.6 Luna | $0.20 | $1.20 | 1.05M | ✅ | OpenAI |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 | 1M | ✅ | |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | 1M | ✅ |
Gemini 2.5 Flash-Lite is the cheapest model with both 1M context and vision support. Only Qwen 3.7 Flash ($0.03/$0.13) and GPT-5 nano ($0.05/$0.40) are cheaper per token, but Qwen is from Alibaba (not a major US provider) and GPT-5 nano lacks vision and has only 128K context.
Which Flash-Lite Tier Should You Use?
📊 2.5: High-Volume Processing
Classification, extraction, moderation, simple Q&A. When cost per token is the primary concern and tasks are straightforward.
🔍 2.5: Document OCR
Image-to-text extraction at scale. Vision support at the lowest price point. Great for receipt processing, form extraction.
🧠 3.1: Moderate Reasoning
Tasks needing better instruction-following and multi-step logic. Summarization, analysis, and content generation.
💻 3.1: Code Assistance
Code completion, explanation, and simple generation. Better reasoning than 2.5 at a moderate premium.
🎯 3.5: Complex Analysis
When you need the strongest Flash-Lite reasoning but still want budget pricing. Research synthesis, nuanced writing.
🤖 3.5: Agentic Workflows
Background agents that need reliable instruction-following. Tool use, multi-step planning, and structured outputs.
Quick Decision Guide
Start with 2.5 Flash-Lite if you're cost-sensitive. It handles most tasks well at $0.10/M input. Upgrade to 3.1 only if you notice quality issues with reasoning or instruction-following. Use 3.5 only when 3.1 isn't enough — the 3x price premium over 2.5 adds up fast at scale.
Alternative: If you don't need vision or Google's ecosystem, Qwen 3.7 Flash ($0.03/$0.13) is 3x cheaper than 2.5 Flash-Lite on input and 3x cheaper on output.
Compare 93 AI Models Side by Side
Gemini Flash-Lite is 3 of 93 models tracked on APIpulse. Compare pricing, context windows, and features across 11 providers.
Frequently Asked Questions
Related Pages
- Qwen 3.7 Flash Pricing — The cheapest AI model at $0.03/$0.13
- GPT-5 nano Pricing — OpenAI's cheapest model at $0.05/$0.40
- GPT-5.6 Luna Pricing — OpenAI's cheapest 1M+ context model
- Google Provider Page — All Gemini models compared
- Top 10 Cheapest LLM APIs — Ranked by input price