DeepSeek V4 Flash API Pricing: 1M Context at $0.22/$0.66

DeepSeek V4 Flash is the cheapest way to get 1M-token context from DeepSeek. At $0.22/$0.66 off-peak, it's ideal for input-heavy tasks like classification, summarization, and RAG.

Updated Aug 21, 2026 · 95 models tracked across 11 providers

TL;DR

⚠️ Peak/Off-Peak Pricing (since Aug 16, 2026): DeepSeek uses time-based pricing. Off-peak (most hours): $0.22/$0.66 per million tokens. Peak (01:00–04:00 and 06:00–10:00 UTC): $0.44/$1.32 — double the off-peak rate. Cached input tokens are $0.0028/M regardless of time.

DeepSeek V4 Flash Pricing Breakdown

Off-peak pricing at different monthly volumes (peak rates are 2× higher):

Monthly Volume Off-Peak Input Off-Peak Output Off-Peak Total (50/50)
1M tokens$0.22$0.66$0.44
10M tokens$2.20$6.60$4.40
100M tokens$22.00$66.00$44.00
1B tokens$220.00$660.00$440.00

At 100M tokens/month off-peak with a 50/50 input/output split, DeepSeek V4 Flash costs $44. Peak hours double these rates. Schedule heavy workloads during off-peak hours to minimize costs.

DeepSeek V4 Family: Flash vs Pro

Model Input $/M Output $/M Context Best For
DeepSeek V4 Flash $0.22 $0.66 1M Input-heavy tasks (classification, summarization, RAG)
DeepSeek V4 Pro $0.66 $1.98 1M Output-heavy tasks (generation, coding)

Decision guide: If your workload is input-heavy (classify, summarize, extract, RAG), choose V4 Flash at $0.22/$0.66. If it's output-heavy (generate, write, code), choose V4 Pro at $0.66/$1.98 — the output savings (3x cheaper) often outweigh the input premium.

DeepSeek V4 Flash vs Other Budget Models

Model Input $/M Output $/M Context Provider
Qwen 3.7 Flash $0.03 $0.13 1M Alibaba
Gemini 2.5 Flash-Lite $0.10 $0.40 1M Google
DeepSeek V4 Flash $0.22 $0.66 1M DeepSeek
GPT-5.4 nano $0.20 $1.25 400K OpenAI
Gemini 2.5 Flash $0.30 $2.50 1M Google
GPT-5.6 Luna $0.20 $1.20 1.05M OpenAI
Claude Haiku 4.5 $1.00 $5.00 200K Anthropic

DeepSeek V4 Flash offers the best balance of price and 1M-context availability. Qwen 3.7 Flash is cheaper but has lower quality. GPT-5.4 nano is comparable on input but 89% more expensive on output. For budget 1M-context work, V4 Flash is hard to beat.

When to Use DeepSeek V4 Flash

📄 Document Classification

Classify thousands of documents by reading their full content. Input-heavy workloads maximize V4 Flash's cheap input pricing at $0.22/M.

📝 Text Summarization

Summarize long articles, reports, or transcripts. 1M context handles the full document; short output keeps costs minimal.

🔍 RAG Pipelines

Retrieve and process large context chunks for question answering. V4 Flash's 1M context and cheap input make it ideal for RAG at scale.

🏷️ Entity Extraction

Extract names, dates, and facts from documents. Short structured output from long input — exactly V4 Flash's sweet spot.

🌐 Translation

Translate documents where input and output lengths are similar. Balanced pricing at $0.22/$0.66 keeps costs predictable.

💬 Chatbots (Short Responses)

Build customer support or FAQ chatbots where responses are brief. The low input cost means cheap conversations at scale.

Real-World Cost: 25K Document Summaries/Month

Suppose you're building a document summarization service that processes 25,000 documents per month, averaging 3,000 input tokens and 500 output tokens per request:

ModelInput CostOutput CostMonthly Total
Qwen 3.7 Flash$2.25$1.63$3.88
DeepSeek V4 Flash$16.50$8.25$24.75
GPT-5.4 nano$15.00$15.63$30.63
Gemini 2.5 Flash$22.50$31.25$53.75
Claude Haiku 4.5$75.00$62.50$137.50

DeepSeek V4 Flash at $24.75/month is 5.5x cheaper than Haiku 4.5 ($137.50) and 46% cheaper than Gemini 2.5 Flash ($53.75) for this input-heavy workload. Qwen 3.7 Flash is even cheaper but may not match V4 Flash's quality for complex summarization.

Compare 95 AI Models Side by Side

DeepSeek V4 Flash is one of 95 models tracked on APIpulse. Compare pricing, context windows, and features across 11 providers.

Frequently Asked Questions

How much does DeepSeek V4 Flash cost?
DeepSeek V4 Flash costs $0.22 per million input tokens and $0.66 per million output tokens during off-peak hours. Peak hours (01:00–04:00 and 06:00–10:00 UTC) double these rates to $0.44/$1.32. Cached input tokens cost $0.0028/M regardless of time.
What is the context window of DeepSeek V4 Flash?
DeepSeek V4 Flash has a 1M token context window. This matches the largest context windows available from Google, OpenAI, and Anthropic, letting you process very long documents or codebases in a single request.
How does DeepSeek V4 Flash compare to DeepSeek V4 Pro?
V4 Flash ($0.22/$0.66) is 3x cheaper on input than V4 Pro ($0.66/$1.98) but 3x more expensive on output. V4 Flash is better for input-heavy tasks like classification, summarization, and RAG. V4 Pro is better for output-heavy tasks like content generation and code writing.
Is DeepSeek V4 Flash cheaper than GPT-5.4 nano?
Yes. DeepSeek V4 Flash ($0.22/$0.66) is slightly more expensive on input than GPT-5.4 nano ($0.20/$1.25) but 47% cheaper on output. V4 Flash also has 1M context vs 400K for nano. For output-heavy tasks or long-context work, V4 Flash is the better value.
When is DeepSeek V4 Flash cheapest?
DeepSeek V4 Flash uses peak/off-peak pricing. Off-peak hours (most of the day) cost $0.22/$0.66 per million tokens. Peak hours (01:00–04:00 and 06:00–10:00 UTC) cost $0.44/$1.32 — double the off-peak rate. Schedule heavy workloads during off-peak hours to minimize costs.
Is DeepSeek V4 Flash available on all platforms?
DeepSeek V4 Flash is available directly through DeepSeek's API and through third-party providers like Together.ai, Fireworks, and others. It's not available on major cloud platforms like AWS Bedrock or Google Vertex AI.

Related Pages