Gemini 3.8 Flash API Pricing: Current and 2027 Rates

Gemini 3.8 Flash currently costs $0.75/M input and $3.75/M output on the paid Standard tier. Those are introductory prices through December 31, 2026—not permanent list prices. Google schedules $1.50/M input and $7.50/M output from January 1, 2027.

Published September 10, 2026 · Pricing checked against official Google Gemini API documentation

Quick answer

Gemini 3.8 Flash price changes

Consumption optionThrough Dec 31, 2026Starting Jan 1, 2027
Standard input$0.75/M$1.50/M
Standard output, including thinking$3.75/M$7.50/M
Batch input$0.375/M$0.75/M
Batch output, including thinking$1.875/M$3.75/M
Flex input$0.375/M$0.75/M
Flex output, including thinking$1.875/M$3.75/M
Priority input$1.35/M$2.70/M
Priority output, including thinking$6.75/M$13.50/M

All figures above are paid-tier USD prices per one million tokens. Free-tier access has quotas and different data-use terms; it is not equivalent to unlimited zero-cost production usage.

Context caching prices

OptionCached input nowCached input from Jan 1, 2027Storage per M token-hours
Standard$0.075/M$0.15/M$0.50 now → $1.00 in 2027
Batch$0.0375/M$0.075/M$0.50 now → $1.00 in 2027
Flex$0.0375/M$0.075/M$0.50 now → $1.00 in 2027
Priority$0.135/M$0.27/M$0.50 now → $1.00 in 2027

How multimodal input is billed

Gemini 3.8 Flash accepts text, image, video, audio and PDF input and returns text. Google publishes one input-token rate across those supported input modalities for this model. That does not make a minute of audio or video equal to one text token: each modality is tokenized, and the resulting token count determines the charge.

No separate output-media price: Gemini 3.8 Flash does not generate audio, images or video. Its output price covers text responses and thinking tokens. Do not apply pricing from Google's separate Live, image-generation or video-generation models to this endpoint.

Example request costs

WorkloadCurrent Standard2027 Standard
10K input + 2K output$0.0150$0.0300
100K input + 20K output$0.1500$0.3000
1M input + 50K output$0.9375$1.8750

These examples assume uncached Standard input and fixed token counts. Real agent costs also depend on thinking output, retries, tool loops, grounding requests and how much context is reused.

When each consumption option fits

Nearby model cost decisions

The current introductory rate makes Gemini 3.8 Flash less expensive per token than many premium reasoning models, but its scheduled 2027 price is double the launch rate. Compare both today's bill and the post-promotion run rate. DeepSeek V4.1 Flash may have lower token prices in published peak/off-peak scenarios, while Claude Fable 5.1 and GPT-6 Astra occupy a much higher price tier. Cost per completed task still depends on quality, tool reliability and retries.

Official sources

Related APIpulse pages