Gemini 3.5 Flash-Lite API Pricing: Google's Thinking Budget Model at $0.30/M Tokens
Gemini 3.5 Flash-Lite is Google's cheapest thinking model — built-in reasoning at $0.30/$2.50 with 1M context, vision, and no hidden thinking token charges.
TL;DR
- Price: $0.30/M input, $2.50/M output — Google's cheapest thinking model. Output price includes thinking tokens
- Context: 1M tokens — one of the largest context windows in the budget tier
- Vision: Yes — text + image input supported. Multimodal reasoning at budget pricing
- Thinking: Yes — built-in step-by-step reasoning. Thinking tokens included in output price, no surprise charges
- Provider: Google AI (ai.google.dev)
- Best for: Tasks needing reasoning at budget prices — data analysis, code generation, document Q&A, nuanced classification
Gemini 3.5 Flash-Lite Pricing Breakdown
At $0.30 per million input tokens and $2.50 per million output tokens, Gemini 3.5 Flash-Lite is Google's cheapest model with thinking capability. The output price includes all reasoning tokens, so you won't get surprise charges from hidden chain-of-thought steps.
| Monthly Volume | Input Cost | Output Cost | Total Cost |
|---|---|---|---|
| 1M tokens | $0.30 | $2.50 | $2.80 |
| 10M tokens | $3.00 | $25.00 | $28.00 |
| 100M tokens | $30.00 | $250.00 | $280.00 |
| 1B tokens | $300.00 | $2,500.00 | $2,800.00 |
At 100M tokens/month, Gemini 3.5 Flash-Lite costs $280. The same volume on Gemini 3.1 Flash-Lite (no thinking) would cost $175, and on Gemini 3 Flash it would cost $350.
How Gemini 3.5 Flash-Lite Compares to Other Budget Models
| Model | Input $/M | Output $/M | Context | Vision | Thinking | Provider |
|---|---|---|---|---|---|---|
| Qwen 3.7 Flash | $0.03 | $0.13 | 1M | ✅ | ❌ | Alibaba |
| DeepSeek V4 Flash | $0.14 | $0.28 | 1M | ❌ | ❌ | DeepSeek |
| GPT-5.4 nano | $0.20 | $1.25 | 400K | ❌ | ❌ | OpenAI |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 | 1M | ✅ | ❌ | |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | 1M | ✅ | ✅ | |
| Gemini 3 Flash | $0.50 | $3.00 | 1M | ✅ | ❌ | |
| GPT-5.4 mini | $0.75 | $4.50 | 400K | ❌ | ❌ | OpenAI |
| Claude Haiku 4.5 | $1.00 | $5.00 | 200K | ❌ | ❌ | Anthropic |
Gemini 3.5 Flash-Lite is the only sub-$1 model with both thinking capability and 1M context. For tasks that need step-by-step reasoning (data analysis, nuanced classification, document Q&A), it offers the best value in the budget tier.
Google Flash-Lite Family: Which One Should You Choose?
| Model | Input $/M | Output $/M | Thinking | Best For |
|---|---|---|---|---|
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | ❌ | Cheapest Google option — simple extraction, classification |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 | ❌ | Budget multimodal — faster, no thinking overhead |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | ✅ | Budget reasoning — step-by-step logic, analysis |
Rule of thumb: Use Gemini 2.5 Flash-Lite for simple, high-volume tasks. Use Gemini 3.1 Flash-Lite when you need better quality but no reasoning. Use Gemini 3.5 Flash-Lite when tasks need step-by-step thinking — the 20% price premium over 3.1 often pays for itself in output quality.
When to Use Gemini 3.5 Flash-Lite
📊 Data Analysis
Analyze datasets, identify patterns, generate insights. Thinking capability handles multi-step analysis without upgrading to a premium model.
💻 Code Generation
Generate and debug code with step-by-step reasoning. Catches edge cases that non-thinking models miss, at budget pricing.
📄 Document Q&A
Answer complex questions about long documents. 1M context handles entire codebases or legal contracts; thinking improves answer quality.
🏷️ Nuanced Classification
Classification tasks that need reasoning — intent detection, sentiment with nuance, multi-label categorization. Thinking helps with ambiguous cases.
🛡️ Customer Support
Handle complex support queries that need multi-step reasoning. Better quality than non-thinking budget models for tricky issues.
🔍 Research Synthesis
Synthesize information from multiple sources, compare arguments, identify gaps. Thinking capability enables deeper analysis at low cost.
When NOT to Use Gemini 3.5 Flash-Lite
Gemini 3.5 Flash-Lite is great for budget reasoning, but consider alternatives when:
- Simple extraction/classification: Gemini 2.5 Flash-Lite ($0.10/$0.40) or Qwen 3.7 Flash ($0.03/$0.13) are 3-8x cheaper when thinking isn't needed
- Pure cost optimization: DeepSeek V4 Flash ($0.14/$0.28) is cheaper for text-only tasks that don't need reasoning
- Complex agentic tasks: Gemini 3.5 Flash ($1.50/$9.00) or Claude Sonnet 5 ($2/$10) have much stronger reasoning
- Creative/nuanced writing: Claude Opus 5 ($5/$25) or GPT-5.4 ($2.50/$15) produce better creative output
- When thinking tokens inflate output: If your task generates lots of reasoning steps, the $2.50/M output can add up — consider Gemini 3.1 Flash-Lite ($0.25/$1.50) for simpler tasks
Real-World Cost Scenario
Let's say you're building a document analysis pipeline that processes 25,000 documents per month, with an average of 3,000 input tokens and 800 output tokens (including thinking) per document:
| Model | Monthly Input | Monthly Output | Total |
|---|---|---|---|
| Qwen 3.7 Flash | $2.25 | $2.60 | $4.85 |
| DeepSeek V4 Flash | $10.50 | $5.60 | $16.10 |
| Gemini 3.1 Flash-Lite | $18.75 | $30.00 | $48.75 |
| Gemini 3.5 Flash-Lite | $22.50 | $50.00 | $72.50 |
| Gemini 3 Flash | $37.50 | $60.00 | $97.50 |
| Claude Haiku 4.5 | $75.00 | $100.00 | $175.00 |
Gemini 3.5 Flash-Lite costs $72.50/month for 25K document analyses — 50% more than Gemini 3.1 Flash-Lite (no thinking) but 26% cheaper than Gemini 3 Flash (also no thinking). The thinking capability often produces better analysis, making the premium worthwhile.
Frequently Asked Questions
Calculate Your Gemini 3.5 Flash-Lite Costs
Enter your token counts and see exactly what you'd pay. Compare against 93 other models instantly.