Gemini 3.1 Flash-Lite API Pricing: 1M Context + Thinking at $0.25/$1.50

Gemini 3.1 Flash-Lite offers 1M-token context, vision, and built-in thinking at a budget price. It bridges the gap between the ultra-cheap 2.5 Flash-Lite and the more capable 3.5 Flash-Lite.

Updated Aug 13, 2026 · 93 models tracked across 11 providers

TL;DR

Thinking on a Budget: Gemini 3.1 Flash-Lite is one of the cheapest models with built-in thinking/reasoning capabilities. At $0.25/$1.50, it's 6x cheaper than Gemini 3.5 Flash-Lite ($0.30/$2.50) while still offering thinking — making it ideal for reasoning tasks where cost matters more than peak quality.

Gemini 3.1 Flash-Lite Pricing Breakdown

Standard pricing at different monthly volumes:

Monthly Volume Input Cost Output Cost Total (70/30 Split)
1M tokens$0.25$1.50$0.63
10M tokens$2.50$15.00$6.25
100M tokens$25.00$150.00$62.50
1B tokens$250.00$1,500.00$625.00

At 100M tokens/month with a 70/30 input/output split, Gemini 3.1 Flash-Lite costs $62.50 — significantly cheaper than the 3.5 Flash-Lite ($105) while still offering thinking capabilities.

Gemini 3.1 Flash-Lite vs Other Budget Models

Model Input $/M Output $/M Context Thinking Provider
Gemini 2.5 Flash-Lite $0.10 $0.40 1M No Google
Gemini 3.1 Flash-Lite $0.25 $1.50 1M Yes Google
Gemini 3.5 Flash-Lite $0.30 $2.50 1M Yes Google
GPT-5.4 nano $0.20 $1.25 400K No OpenAI
Qwen 3.7 Flash $0.03 $0.13 1M No Alibaba
DeepSeek V4 Flash $0.14 $0.28 1M No DeepSeek

Gemini 3.1 Flash-Lite is the cheapest model with both 1M context and thinking. Qwen 3.7 Flash is cheaper but lacks thinking; Gemini 3.5 Flash-Lite has thinking but costs more.

Google Flash-Lite Family: Which One to Choose?

Model Input $/M Output $/M Context Thinking Best For
Gemini 2.5 Flash-Lite $0.10 $0.40 1M No High-volume, simple tasks
Gemini 3.1 Flash-Lite $0.25 $1.50 1M Yes Reasoning on a budget
Gemini 3.5 Flash-Lite $0.30 $2.50 1M Yes Best thinking quality

Decision guide: Need the absolute cheapest? → 2.5 Flash-Lite. Need thinking on a budget? → 3.1 Flash-Lite. Need the best thinking quality? → 3.5 Flash-Lite.

When to Use Gemini 3.1 Flash-Lite

🧠 Reasoning Tasks

Tasks that need step-by-step thinking — classification, extraction, analysis — at budget pricing. The built-in thinking improves accuracy without premium model costs.

📄 Document Analysis

Process long documents with reasoning about their content. 1M context handles entire books; thinking improves comprehension and summary quality.

👁️ Multimodal Processing

Analyze images, videos, or audio with reasoning capabilities. Vision inputs cost the same as text — no premium for multimodal.

🔍 RAG Applications

Large context + thinking means better reasoning over retrieved documents. Improved answer quality without the cost of premium models.

📊 Data Classification

Classify and categorize data with reasoning about edge cases. Thinking helps handle ambiguous inputs that simpler models get wrong.

💬 Customer Support

Handle complex support queries that need reasoning about context. Thinking improves response quality for nuanced customer issues.

Real-World Cost: Processing 25,000 Documents/Month

Suppose you're building a document analysis service that processes 25,000 documents per month, averaging 3,000 input tokens and generating 800 output tokens per document:

ModelInput CostOutput CostMonthly Total
Gemini 2.5 Flash-Lite$7.50$8.00$15.50
Gemini 3.1 Flash-Lite$18.75$30.00$48.75
Gemini 3.5 Flash-Lite$22.50$50.00$72.50
GPT-5.4 nano$15.00$25.00$40.00
Qwen 3.7 Flash$2.25$2.60$4.85

At $48.75/month, Gemini 3.1 Flash-Lite costs 3x more than 2.5 Flash-Lite but delivers thinking capabilities. For tasks where reasoning accuracy matters, the improved quality often justifies the cost difference.

Compare 93 AI Models Side by Side

Gemini 3.1 Flash-Lite is one of 93 models tracked on APIpulse. Compare pricing, context windows, and features across 11 providers.

Frequently Asked Questions

How much does Gemini 3.1 Flash-Lite cost?
Gemini 3.1 Flash-Lite costs $0.25 per million input tokens and $1.50 per million output tokens. This places it between the cheaper 2.5 Flash-Lite ($0.10/$0.40) and the more capable 3.5 Flash-Lite ($0.30/$2.50).
What is the context window of Gemini 3.1 Flash-Lite?
Gemini 3.1 Flash-Lite has a 1M token context window, matching every other model in Google's Flash-Lite family. This lets you process long documents, codebases, or conversation histories in a single request.
Does Gemini 3.1 Flash-Lite support vision and thinking?
Yes. Gemini 3.1 Flash-Lite supports multimodal inputs (text, images, video, audio) and has built-in thinking/reasoning capabilities. This makes it more capable than the 2.5 Flash-Lite, which lacks thinking.
How does Gemini 3.1 Flash-Lite compare to 2.5 Flash-Lite?
3.1 Flash-Lite ($0.25/$1.50) is 2.5x more expensive than 2.5 Flash-Lite ($0.10/$0.40) on input and 3.75x on output. The extra cost buys thinking capabilities and improved reasoning. For simple tasks, 2.5 Flash-Lite is cheaper; for tasks needing reasoning, 3.1 Flash-Lite is the better choice.
Is Gemini 3.1 Flash-Lite being deprecated?
No. Gemini 3.1 Flash-Lite is actively available on Google AI Studio and Vertex AI. Google has released newer Flash-Lite models (3.5) but 3.1 remains available and is not scheduled for deprecation.

Related Pages