Gemini 3.8 Flash API Pricing: Current and 2027 Rates
Gemini 3.8 Flash currently costs $0.75/M input and $3.75/M output on the paid Standard tier. Those are introductory prices through December 31, 2026—not permanent list prices. Google schedules $1.50/M input and $7.50/M output from January 1, 2027.
- Current Standard: $0.75/M input · $3.75/M output through December 31, 2026.
- Scheduled January 1, 2027: $1.50/M input · $7.50/M output.
- Current Batch and Flex: $0.375/M input · $1.875/M output.
- Current Priority: $1.35/M input · $6.75/M output.
- Model ID:
gemini-3.8-flash; status: GA.
Gemini 3.8 Flash price changes
| Consumption option | Through Dec 31, 2026 | Starting Jan 1, 2027 |
|---|---|---|
| Standard input | $0.75/M | $1.50/M |
| Standard output, including thinking | $3.75/M | $7.50/M |
| Batch input | $0.375/M | $0.75/M |
| Batch output, including thinking | $1.875/M | $3.75/M |
| Flex input | $0.375/M | $0.75/M |
| Flex output, including thinking | $1.875/M | $3.75/M |
| Priority input | $1.35/M | $2.70/M |
| Priority output, including thinking | $6.75/M | $13.50/M |
All figures above are paid-tier USD prices per one million tokens. Free-tier access has quotas and different data-use terms; it is not equivalent to unlimited zero-cost production usage.
Context caching prices
| Option | Cached input now | Cached input from Jan 1, 2027 | Storage per M token-hours |
|---|---|---|---|
| Standard | $0.075/M | $0.15/M | $0.50 now → $1.00 in 2027 |
| Batch | $0.0375/M | $0.075/M | $0.50 now → $1.00 in 2027 |
| Flex | $0.0375/M | $0.075/M | $0.50 now → $1.00 in 2027 |
| Priority | $0.135/M | $0.27/M | $0.50 now → $1.00 in 2027 |
How multimodal input is billed
Gemini 3.8 Flash accepts text, image, video, audio and PDF input and returns text. Google publishes one input-token rate across those supported input modalities for this model. That does not make a minute of audio or video equal to one text token: each modality is tokenized, and the resulting token count determines the charge.
Example request costs
| Workload | Current Standard | 2027 Standard |
|---|---|---|
| 10K input + 2K output | $0.0150 | $0.0300 |
| 100K input + 20K output | $0.1500 | $0.3000 |
| 1M input + 50K output | $0.9375 | $1.8750 |
These examples assume uncached Standard input and fixed token counts. Real agent costs also depend on thinking output, retries, tool loops, grounding requests and how much context is reused.
When each consumption option fits
- Standard: ordinary interactive API traffic and the clearest baseline for budgeting.
- Batch: asynchronous evaluations or bulk processing that can wait.
- Flex: lower-cost traffic that can accept reduced availability and higher latency.
- Priority: workloads paying more for prioritized serving behavior; do not choose it merely because it appears in the pricing table.
Nearby model cost decisions
The current introductory rate makes Gemini 3.8 Flash less expensive per token than many premium reasoning models, but its scheduled 2027 price is double the launch rate. Compare both today's bill and the post-promotion run rate. DeepSeek V4.1 Flash may have lower token prices in published peak/off-peak scenarios, while Claude Fable 5.1 and GPT-6 Astra occupy a much higher price tier. Cost per completed task still depends on quality, tool reliability and retries.
Official sources
- Google Gemini Developer API pricing
- Google Gemini 3.8 Flash model documentation
- Google Gemini 3.8 Flash release and migration guide