Gemini 3.1 Flash-Lite API Pricing: 1M Context + Thinking at $0.25/$1.50
Gemini 3.1 Flash-Lite offers 1M-token context, vision, and built-in thinking at a budget price. It bridges the gap between the ultra-cheap 2.5 Flash-Lite and the more capable 3.5 Flash-Lite.
TL;DR
- Price: $0.25/M input, $1.50/M output — mid-tier in Google's Flash-Lite family
- Context: 1M tokens — process entire books or codebases in one request
- Capabilities: Vision, multimodal input, and built-in thinking/reasoning
- Provider: Google AI Studio / Vertex AI
- Best for: Reasoning tasks on a budget, document analysis, multimodal processing
- Trade-off: 2.5x pricier than 2.5 Flash-Lite — choose based on whether you need thinking
Gemini 3.1 Flash-Lite Pricing Breakdown
Standard pricing at different monthly volumes:
| Monthly Volume | Input Cost | Output Cost | Total (70/30 Split) |
|---|---|---|---|
| 1M tokens | $0.25 | $1.50 | $0.63 |
| 10M tokens | $2.50 | $15.00 | $6.25 |
| 100M tokens | $25.00 | $150.00 | $62.50 |
| 1B tokens | $250.00 | $1,500.00 | $625.00 |
At 100M tokens/month with a 70/30 input/output split, Gemini 3.1 Flash-Lite costs $62.50 — significantly cheaper than the 3.5 Flash-Lite ($105) while still offering thinking capabilities.
Gemini 3.1 Flash-Lite vs Other Budget Models
| Model | Input $/M | Output $/M | Context | Thinking | Provider |
|---|---|---|---|---|---|
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | 1M | No | |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 | 1M | Yes | |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | 1M | Yes | |
| GPT-5.4 nano | $0.20 | $1.25 | 400K | No | OpenAI |
| Qwen 3.7 Flash | $0.03 | $0.13 | 1M | No | Alibaba |
| DeepSeek V4 Flash | $0.14 | $0.28 | 1M | No | DeepSeek |
Gemini 3.1 Flash-Lite is the cheapest model with both 1M context and thinking. Qwen 3.7 Flash is cheaper but lacks thinking; Gemini 3.5 Flash-Lite has thinking but costs more.
Google Flash-Lite Family: Which One to Choose?
| Model | Input $/M | Output $/M | Context | Thinking | Best For |
|---|---|---|---|---|---|
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | 1M | No | High-volume, simple tasks |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 | 1M | Yes | Reasoning on a budget |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | 1M | Yes | Best thinking quality |
Decision guide: Need the absolute cheapest? → 2.5 Flash-Lite. Need thinking on a budget? → 3.1 Flash-Lite. Need the best thinking quality? → 3.5 Flash-Lite.
When to Use Gemini 3.1 Flash-Lite
🧠 Reasoning Tasks
Tasks that need step-by-step thinking — classification, extraction, analysis — at budget pricing. The built-in thinking improves accuracy without premium model costs.
📄 Document Analysis
Process long documents with reasoning about their content. 1M context handles entire books; thinking improves comprehension and summary quality.
👁️ Multimodal Processing
Analyze images, videos, or audio with reasoning capabilities. Vision inputs cost the same as text — no premium for multimodal.
🔍 RAG Applications
Large context + thinking means better reasoning over retrieved documents. Improved answer quality without the cost of premium models.
📊 Data Classification
Classify and categorize data with reasoning about edge cases. Thinking helps handle ambiguous inputs that simpler models get wrong.
💬 Customer Support
Handle complex support queries that need reasoning about context. Thinking improves response quality for nuanced customer issues.
Real-World Cost: Processing 25,000 Documents/Month
Suppose you're building a document analysis service that processes 25,000 documents per month, averaging 3,000 input tokens and generating 800 output tokens per document:
| Model | Input Cost | Output Cost | Monthly Total |
|---|---|---|---|
| Gemini 2.5 Flash-Lite | $7.50 | $8.00 | $15.50 |
| Gemini 3.1 Flash-Lite | $18.75 | $30.00 | $48.75 |
| Gemini 3.5 Flash-Lite | $22.50 | $50.00 | $72.50 |
| GPT-5.4 nano | $15.00 | $25.00 | $40.00 |
| Qwen 3.7 Flash | $2.25 | $2.60 | $4.85 |
At $48.75/month, Gemini 3.1 Flash-Lite costs 3x more than 2.5 Flash-Lite but delivers thinking capabilities. For tasks where reasoning accuracy matters, the improved quality often justifies the cost difference.
Compare 93 AI Models Side by Side
Gemini 3.1 Flash-Lite is one of 93 models tracked on APIpulse. Compare pricing, context windows, and features across 11 providers.
Frequently Asked Questions
Related Pages
- Gemini 2.5 Flash-Lite Pricing — Cheaper at $0.10/$0.40 but no thinking
- Gemini 3.5 Flash-Lite Pricing — Better thinking at $0.30/$2.50
- All Google Flash-Lite Models — Compare all Flash-Lite tiers
- Qwen 3.7 Flash Pricing — Cheapest model at $0.03/$0.13
- Full Model Rankings — All 93 models ranked by price