Gemini Flash-Lite API Pricing: All 3 Google Budget Tiers Compared

Google offers three Flash-Lite models from $0.10/M to $0.30/M — all with 1M context and vision support. Here's which tier to pick.

Updated Aug 7, 2026 · 93 models tracked across 11 providers

TL;DR

The 3 Gemini Flash-Lite Tiers

Gemini 3.1 Flash-Lite

$0.25/M input

$1.50/M output · 1M context

  • Better reasoning
  • Vision support
  • Good balance

Gemini 3.5 Flash-Lite

$0.30/M input

$2.50/M output · 1M context

  • Strongest capabilities
  • Vision support
  • Latest architecture

Flash-Lite Pricing Comparison

Model Input $/M Output $/M Context Vision Generation
Gemini 2.5 Flash-Lite $0.10 $0.40 1M 2.5
Gemini 3.1 Flash-Lite $0.25 $1.50 1M 3.1
Gemini 3.5 Flash-Lite $0.30 $2.50 1M 3.5

The price gap between tiers is significant: 3.1 is 2.5x more expensive on input than 2.5, and 3.5 is 3x more. For most budget workloads, 2.5 Flash-Lite offers the best cost-per-token.

Monthly Cost at Different Volumes

Monthly Volume 2.5 Flash-Lite 3.1 Flash-Lite 3.5 Flash-Lite
1M tokens$0.50$1.75$2.80
10M tokens$5.00$17.50$28.00
100M tokens$50.00$175.00$280.00
1B tokens$500.00$1,750.00$2,800.00

At 100M tokens/month, choosing 2.5 over 3.5 saves $230/month. The cost difference is negligible for small volumes but compounds fast at scale.

Flash-Lite vs Other Budget Models

Model Input $/M Output $/M Context Vision Provider
Qwen 3.7 Flash $0.03 $0.13 1M Alibaba
GPT-5 nano $0.05 $0.40 128K OpenAI
Gemini 2.5 Flash-Lite $0.10 $0.40 1M Google
DeepSeek V4 Flash $0.14 $0.28 1M DeepSeek
GPT-5.6 Luna $0.20 $1.20 1.05M OpenAI
Gemini 3.1 Flash-Lite $0.25 $1.50 1M Google
Gemini 3.5 Flash-Lite $0.30 $2.50 1M Google

Gemini 2.5 Flash-Lite is the cheapest model with both 1M context and vision support. Only Qwen 3.7 Flash ($0.03/$0.13) and GPT-5 nano ($0.05/$0.40) are cheaper per token, but Qwen is from Alibaba (not a major US provider) and GPT-5 nano lacks vision and has only 128K context.

Which Flash-Lite Tier Should You Use?

📊 2.5: High-Volume Processing

Classification, extraction, moderation, simple Q&A. When cost per token is the primary concern and tasks are straightforward.

🔍 2.5: Document OCR

Image-to-text extraction at scale. Vision support at the lowest price point. Great for receipt processing, form extraction.

🧠 3.1: Moderate Reasoning

Tasks needing better instruction-following and multi-step logic. Summarization, analysis, and content generation.

💻 3.1: Code Assistance

Code completion, explanation, and simple generation. Better reasoning than 2.5 at a moderate premium.

🎯 3.5: Complex Analysis

When you need the strongest Flash-Lite reasoning but still want budget pricing. Research synthesis, nuanced writing.

🤖 3.5: Agentic Workflows

Background agents that need reliable instruction-following. Tool use, multi-step planning, and structured outputs.

Quick Decision Guide

Start with 2.5 Flash-Lite if you're cost-sensitive. It handles most tasks well at $0.10/M input. Upgrade to 3.1 only if you notice quality issues with reasoning or instruction-following. Use 3.5 only when 3.1 isn't enough — the 3x price premium over 2.5 adds up fast at scale.

Alternative: If you don't need vision or Google's ecosystem, Qwen 3.7 Flash ($0.03/$0.13) is 3x cheaper than 2.5 Flash-Lite on input and 3x cheaper on output.

Compare 93 AI Models Side by Side

Gemini Flash-Lite is 3 of 93 models tracked on APIpulse. Compare pricing, context windows, and features across 11 providers.

Frequently Asked Questions

How much does Gemini Flash-Lite cost?
Google offers 3 Flash-Lite tiers: Gemini 2.5 Flash-Lite at $0.10/$0.40 per million tokens, Gemini 3.1 Flash-Lite at $0.25/$1.50, and Gemini 3.5 Flash-Lite at $0.30/$2.50. All three have 1M token context windows and support vision.
Which Gemini Flash-Lite should I use?
Use 2.5 Flash-Lite ($0.10/$0.40) for the cheapest option — ideal for classification, extraction, and simple tasks. Use 3.1 Flash-Lite ($0.25/$1.50) for better reasoning at a moderate premium. Use 3.5 Flash-Lite ($0.30/$2.50) when you need the strongest Flash-Lite capabilities but still want budget pricing.
Is Gemini Flash-Lite cheaper than GPT-5 nano?
Gemini 2.5 Flash-Lite ($0.10/$0.40) is 2x more expensive on input than GPT-5 nano ($0.05/$0.40) but offers 1M context vs 128K, plus vision support. Qwen 3.7 Flash ($0.03/$0.13) is cheaper than both.
Do all Gemini Flash-Lite models support vision?
Yes. All three Gemini Flash-Lite models (2.5, 3.1, and 3.5) support image input alongside text. This makes them some of the cheapest vision-capable models available.
What is the context window of Gemini Flash-Lite?
All three Gemini Flash-Lite models have 1 million token context windows. This is the same as Gemini Flash and most other Google models, and significantly larger than OpenAI's budget models (128K for GPT-5 nano).

Related Pages