Gemini 2.5 Flash API Pricing: 1M Context at $0.30/$2.50

Gemini 2.5 Flash offers Google's 1M-token context window with vision support at a budget price. Batch API drops costs to $0.15/$1.25 — half the standard rate.

Updated Aug 13, 2026 · 93 models tracked across 11 providers

TL;DR

Batch API Saves 50%: Google's batch API cuts Gemini 2.5 Flash pricing in half — $0.15/M input, $1.25/M output. Batch jobs process asynchronously (typically within 24 hours). Use this for any non-real-time workload: data processing, content generation, analysis.

Gemini 2.5 Flash Pricing Breakdown

Standard and batch API pricing at different monthly volumes:

Monthly Volume Standard Input Standard Output Batch Input Batch Output
1M tokens$0.30$2.50$0.15$1.25
10M tokens$3.00$25.00$1.50$12.50
100M tokens$30.00$250.00$15.00$125.00
1B tokens$300.00$2,500.00$150.00$1,250.00

At 100M tokens/month with a 70/30 input/output split, standard API costs $105/month. Batch API drops this to $52.50 — making it competitive with many smaller-context models.

Gemini 2.5 Flash vs Other 1M-Context Models

Model Input $/M Output $/M Context Vision Provider
GPT-5.6 Luna $0.20 $1.20 1.05M Yes OpenAI
GPT-4.1 nano $0.10 $0.40 1M Yes OpenAI
Gemini 2.5 Flash-Lite $0.10 $0.40 1M Yes Google
Gemini 2.5 Flash $0.30 $2.50 1M Yes Google
Gemini 3 Flash $0.50 $3.00 1M Yes Google
Claude Sonnet 5 $2.00 $10.00 1M Yes Anthropic

Gemini 2.5 Flash sits in the middle of the 1M-context price range. It's more expensive than Flash-Lite or GPT-4.1 nano, but offers stronger reasoning and better multimodal quality. For tasks that need both long context and quality, it's a strong value.

Google Flash Family: Which One to Choose?

Model Input $/M Output $/M Context Best For
Gemini 2.5 Flash-Lite $0.10 $0.40 1M High-volume, simple tasks
Gemini 2.5 Flash $0.30 $2.50 1M Balanced quality/cost
Gemini 3 Flash $0.50 $3.00 1M Latest capabilities
Gemini 3.5 Flash $1.50 $9.00 1M Complex reasoning
Gemini 3.6 Flash $1.50 $7.50 1M Newest, best quality

Gemini 2.5 Flash is the sweet spot for budget-conscious users who need more quality than Flash-Lite but don't want to pay for the newer 3.x models. It's the most cost-effective Google model with both 1M context and strong reasoning.

When to Use Gemini 2.5 Flash

📄 Long Documents

Process entire books, legal contracts, or codebases in a single request. The 1M context window handles what smaller models can't.

👁️ Multimodal Analysis

Analyze images, videos, or audio alongside text. Vision inputs cost the same as text — no premium for multimodal.

📦 Batch Processing

Use the batch API for 50% off. Perfect for data processing, content generation, or analysis that doesn't need real-time results.

🔍 RAG Applications

Large context windows mean you can stuff more retrieved documents into each request, improving answer quality without extra API calls.

💻 Code Analysis

Feed entire repositories for code review, documentation generation, or refactoring suggestions. 1M tokens fits most codebases.

📊 Data Extraction

Extract structured data from long unstructured documents. Strong reasoning at budget pricing makes it ideal for high-volume extraction.

Real-World Cost: Processing 10,000 Documents/Month

Suppose you're building a document analysis service that processes 10,000 documents per month, averaging 5,000 input tokens and generating 1,000 output tokens per document:

ModelInput CostOutput CostMonthly Total
Gemini 2.5 Flash (Batch)$7.50$12.50$20.00
Gemini 2.5 Flash (Standard)$15.00$25.00$40.00
Gemini 3 Flash$25.00$30.00$55.00
GPT-5.6 Luna$10.00$12.00$22.00
Claude Sonnet 5$100.00$100.00$200.00

With batch API, Gemini 2.5 Flash at $20/month is one of the cheapest 1M-context options. Even standard pricing ($40) is reasonable for the quality and context window you get.

Compare 93 AI Models Side by Side

Gemini 2.5 Flash is one of 93 models tracked on APIpulse. Compare pricing, context windows, and features across 11 providers.

Frequently Asked Questions

How much does Gemini 2.5 Flash cost?
Gemini 2.5 Flash costs $0.30 per million input tokens and $2.50 per million output tokens. With batch API, prices drop to $0.15/$1.25 — a 50% discount.
What is the context window of Gemini 2.5 Flash?
Gemini 2.5 Flash has a 1M token context window. This is among the largest available, letting you process long documents, codebases, or conversation histories in a single request.
Does Gemini 2.5 Flash support vision?
Yes. Gemini 2.5 Flash is a multimodal model that accepts text, images, video, and audio inputs. Vision and multimodal inputs are priced at the same $0.30/M rate as text input.
Is Gemini 2.5 Flash cheaper than GPT-5.4 nano?
GPT-5.4 nano ($0.20/$1.25) is cheaper on both input and output. However, Gemini 2.5 Flash offers 1M context (vs 400K), vision/multimodal support, and batch API at $0.15/$1.25. For multimodal or long-context tasks, Gemini 2.5 Flash is the better value.
What is the Gemini 2.5 Flash batch API discount?
The batch API offers a 50% discount: $0.15/M input and $1.25/M output (vs $0.30/$2.50 standard). Batch jobs process asynchronously and complete within 24 hours. This is ideal for non-real-time workloads like data processing or content generation.
Is Gemini 2.5 Flash being deprecated?
No. Gemini 2.5 Flash is still actively available. Google has released newer Flash models (3.0, 3.1, 3.5, 3.6) but 2.5 Flash remains available and is not scheduled for deprecation. It's a safe choice for production workloads.

Related Pages