DeepSeek V4 Flash API Pricing: Retired — Served as V4.1 Flash

DeepSeek V4 Flash has been retired. The legacy deepseek-v4-flash name still works, but requests run on DeepSeek-V4.1-Flash and are billed at $0.15/$0.60 off-peak — cheaper than V4 Flash's historical $0.22/$0.66.

Updated Sep 25, 2026 · 120 models tracked across 16 providers

✅ Status update (Sep 25, 2026): V4 Flash and the vision-experimental variant are retired per DeepSeek's documentation. The names deepseek-v4-flash and deepseek-v4-flash-vision-exp are still accepted and served by DeepSeek-V4.1-Flash (requested as deepseek-flash) at the V4.1 Flash price: $0.15 input / $0.60 output off-peak ($0.30/$1.20 peak), 1M context, 384K max output, vision supported. The pricing tables below reflect V4 Flash's pre-retirement rates and are kept for historical reference. Current numbers live in our DeepSeek pricing guide.

TL;DR

⚠️ Peak/Off-Peak Pricing (since Aug 16, 2026): DeepSeek uses time-based pricing. Off-peak (most hours): V4 Flash billed $0.22/$0.66 per million tokens historically; V4.1 Flash bills $0.15/$0.60. Peak (01:00–04:00 and 06:00–10:00 UTC): double the off-peak rate. Since Aug 23: weekends (Sat/Sun Beijing time) are all-day off-peak.

DeepSeek V4 Flash Pricing Breakdown

Off-peak pricing at different monthly volumes (peak rates are 2× higher):

Monthly Volume Off-Peak Input Off-Peak Output Off-Peak Total (50/50)
1M tokens$0.22$0.66$0.44
10M tokens$2.20$6.60$4.40
100M tokens$22.00$66.00$44.00
1B tokens$220.00$660.00$440.00

At 100M tokens/month off-peak with a 50/50 input/output split, DeepSeek V4 Flash costs $44. Peak hours double these rates. Schedule heavy workloads during off-peak hours to minimize costs.

DeepSeek V4 Family: Flash vs Pro

Model Input $/M Output $/M Context Best For
DeepSeek V4 Flash $0.22 $0.66 1M Input-heavy tasks (classification, summarization, RAG)
DeepSeek V4 Pro $0.66 $1.98 1M Output-heavy tasks (generation, coding)

Decision guide: If your workload is input-heavy (classify, summarize, extract, RAG), choose V4 Flash at $0.22/$0.66. If it's output-heavy (generate, write, code), choose V4 Pro at $0.66/$1.98 — the output savings (3x cheaper) often outweigh the input premium.

DeepSeek V4 Flash vs Other Budget Models

Model Input $/M Output $/M Context Provider
Qwen 3.7 Flash $0.03 $0.13 1M Alibaba
Gemini 2.5 Flash-Lite $0.10 $0.40 1M Google
DeepSeek V4 Flash $0.22 $0.66 1M DeepSeek
GPT-5.4 nano $0.20 $1.25 400K OpenAI
Gemini 2.5 Flash $0.30 $2.50 1M Google
GPT-5.6 Luna $0.20 $1.20 1.05M OpenAI
Claude Haiku 4.5 $1.00 $5.00 200K Anthropic

DeepSeek V4 Flash offers the best balance of price and 1M-context availability. Qwen 3.7 Flash is cheaper but has lower quality. GPT-5.4 nano is comparable on input but 89% more expensive on output. For budget 1M-context work, V4 Flash is hard to beat.

When to Use DeepSeek V4 Flash

📄 Document Classification

Classify thousands of documents by reading their full content. Input-heavy workloads maximize V4 Flash's cheap input pricing at $0.22/M.

📝 Text Summarization

Summarize long articles, reports, or transcripts. 1M context handles the full document; short output keeps costs minimal.

🔍 RAG Pipelines

Retrieve and process large context chunks for question answering. V4 Flash's 1M context and cheap input make it ideal for RAG at scale.

🏷️ Entity Extraction

Extract names, dates, and facts from documents. Short structured output from long input — exactly V4 Flash's sweet spot.

🌐 Translation

Translate documents where input and output lengths are similar. Balanced pricing at $0.22/$0.66 keeps costs predictable.

💬 Chatbots (Short Responses)

Build customer support or FAQ chatbots where responses are brief. The low input cost means cheap conversations at scale.

Real-World Cost: 25K Document Summaries/Month

Suppose you're building a document summarization service that processes 25,000 documents per month, averaging 3,000 input tokens and 500 output tokens per request:

ModelInput CostOutput CostMonthly Total
Qwen 3.7 Flash$2.25$1.63$3.88
DeepSeek V4 Flash$16.50$8.25$24.75
GPT-5.4 nano$15.00$15.63$30.63
Gemini 2.5 Flash$22.50$31.25$53.75
Claude Haiku 4.5$75.00$62.50$137.50

DeepSeek V4 Flash at $24.75/month is 5.5x cheaper than Haiku 4.5 ($137.50) and 46% cheaper than Gemini 2.5 Flash ($53.75) for this input-heavy workload. Qwen 3.7 Flash is even cheaper but may not match V4 Flash's quality for complex summarization.

Compare 95 AI Models Side by Side

DeepSeek V4 Flash is one of 120 models tracked on APIpulse. Compare pricing, context windows, and features across 16 providers.

Frequently Asked Questions

How much does DeepSeek V4 Flash cost?
DeepSeek V4 Flash is retired. Legacy deepseek-v4-flash requests are now served by DeepSeek-V4.1-Flash and billed at $0.15 per million input tokens and $0.60 per million output tokens off-peak ($0.30/$1.20 at peak). Before retirement V4 Flash charged $0.22/$0.66 off-peak, doubling to $0.44/$1.32 during peak hours (01:00–04:00 and 06:00–10:00 UTC). Since Aug 23, 2026, weekends (Sat/Sun Beijing time) are all-day off-peak.
What is the context window of DeepSeek V4 Flash?
DeepSeek V4 Flash has a 1M token context window. This matches the largest context windows available from Google, OpenAI, and Anthropic, letting you process very long documents or codebases in a single request.
How does DeepSeek V4 Flash compare to DeepSeek V4 Pro?
Legacy V4 Flash requests now bill at V4.1 Flash rates ($0.15/$0.60 off-peak), which are 4.4x cheaper on input than V4 Pro ($0.66/$1.98) but 3.3x more expensive on output. The Flash line is better for input-heavy tasks like classification, summarization, and RAG. V4 Pro is better for output-heavy tasks like content generation and code writing.
Is DeepSeek V4 Flash cheaper than GPT-5.4 nano?
Yes, and the gap widened after V4 Flash was retired: DeepSeek-V4.1-Flash bills legacy V4 Flash requests at $0.15/$0.60 off-peak versus GPT-5.4 nano at $0.20/$1.25 — 25% cheaper on input and 52% cheaper on output. V4.1 Flash also has 1M context vs 400K for nano. For output-heavy tasks or long-context work, the DeepSeek Flash line is the better value.
When is DeepSeek V4 Flash cheapest?
The DeepSeek Flash line uses peak/off-peak pricing. Off-peak hours (most of the day) cost $0.15/$0.60 per million tokens on DeepSeek-V4.1-Flash, which serves legacy V4 Flash requests. Peak hours (01:00–04:00 and 06:00–10:00 UTC) cost $0.30/$1.20 — double the off-peak rate. Since Aug 23, 2026, weekends (Sat/Sun Beijing time) are all-day off-peak. Schedule heavy workloads during off-peak hours and weekends to minimize costs.
Is DeepSeek V4 Flash available on all platforms?
DeepSeek V4 Flash is available directly through DeepSeek's API and through third-party providers like Together.ai, Fireworks, and others. It's not available on major cloud platforms like AWS Bedrock or Google Vertex AI.

Related Pages