DeepSeek V4 Flash API Pricing: 1M Context at $0.22/$0.66
DeepSeek V4 Flash is the cheapest way to get 1M-token context from DeepSeek. At $0.22/$0.66 off-peak, it's ideal for input-heavy tasks like classification, summarization, and RAG.
TL;DR
- Price: $0.22/M input, $0.66/M output (off-peak) — cheapest DeepSeek model with 1M context
- Context: 1M tokens — process entire books or codebases in one request
- Provider: DeepSeek API (+ Together.ai, Fireworks, others)
- Best for: Input-heavy tasks — classification, summarization, RAG, extraction
- Trade-off: Output is 3x pricier than V4 Pro — choose V4 Pro for generation-heavy workloads
DeepSeek V4 Flash Pricing Breakdown
Off-peak pricing at different monthly volumes (peak rates are 2× higher):
| Monthly Volume | Off-Peak Input | Off-Peak Output | Off-Peak Total (50/50) |
|---|---|---|---|
| 1M tokens | $0.22 | $0.66 | $0.44 |
| 10M tokens | $2.20 | $6.60 | $4.40 |
| 100M tokens | $22.00 | $66.00 | $44.00 |
| 1B tokens | $220.00 | $660.00 | $440.00 |
At 100M tokens/month off-peak with a 50/50 input/output split, DeepSeek V4 Flash costs $44. Peak hours double these rates. Schedule heavy workloads during off-peak hours to minimize costs.
DeepSeek V4 Family: Flash vs Pro
| Model | Input $/M | Output $/M | Context | Best For |
|---|---|---|---|---|
| DeepSeek V4 Flash | $0.22 | $0.66 | 1M | Input-heavy tasks (classification, summarization, RAG) |
| DeepSeek V4 Pro | $0.66 | $1.98 | 1M | Output-heavy tasks (generation, coding) |
Decision guide: If your workload is input-heavy (classify, summarize, extract, RAG), choose V4 Flash at $0.22/$0.66. If it's output-heavy (generate, write, code), choose V4 Pro at $0.66/$1.98 — the output savings (3x cheaper) often outweigh the input premium.
DeepSeek V4 Flash vs Other Budget Models
| Model | Input $/M | Output $/M | Context | Provider |
|---|---|---|---|---|
| Qwen 3.7 Flash | $0.03 | $0.13 | 1M | Alibaba |
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | 1M | |
| DeepSeek V4 Flash | $0.22 | $0.66 | 1M | DeepSeek |
| GPT-5.4 nano | $0.20 | $1.25 | 400K | OpenAI |
| Gemini 2.5 Flash | $0.30 | $2.50 | 1M | |
| GPT-5.6 Luna | $0.20 | $1.20 | 1.05M | OpenAI |
| Claude Haiku 4.5 | $1.00 | $5.00 | 200K | Anthropic |
DeepSeek V4 Flash offers the best balance of price and 1M-context availability. Qwen 3.7 Flash is cheaper but has lower quality. GPT-5.4 nano is comparable on input but 89% more expensive on output. For budget 1M-context work, V4 Flash is hard to beat.
When to Use DeepSeek V4 Flash
📄 Document Classification
Classify thousands of documents by reading their full content. Input-heavy workloads maximize V4 Flash's cheap input pricing at $0.22/M.
📝 Text Summarization
Summarize long articles, reports, or transcripts. 1M context handles the full document; short output keeps costs minimal.
🔍 RAG Pipelines
Retrieve and process large context chunks for question answering. V4 Flash's 1M context and cheap input make it ideal for RAG at scale.
🏷️ Entity Extraction
Extract names, dates, and facts from documents. Short structured output from long input — exactly V4 Flash's sweet spot.
🌐 Translation
Translate documents where input and output lengths are similar. Balanced pricing at $0.22/$0.66 keeps costs predictable.
💬 Chatbots (Short Responses)
Build customer support or FAQ chatbots where responses are brief. The low input cost means cheap conversations at scale.
Real-World Cost: 25K Document Summaries/Month
Suppose you're building a document summarization service that processes 25,000 documents per month, averaging 3,000 input tokens and 500 output tokens per request:
| Model | Input Cost | Output Cost | Monthly Total |
|---|---|---|---|
| Qwen 3.7 Flash | $2.25 | $1.63 | $3.88 |
| DeepSeek V4 Flash | $16.50 | $8.25 | $24.75 |
| GPT-5.4 nano | $15.00 | $15.63 | $30.63 |
| Gemini 2.5 Flash | $22.50 | $31.25 | $53.75 |
| Claude Haiku 4.5 | $75.00 | $62.50 | $137.50 |
DeepSeek V4 Flash at $24.75/month is 5.5x cheaper than Haiku 4.5 ($137.50) and 46% cheaper than Gemini 2.5 Flash ($53.75) for this input-heavy workload. Qwen 3.7 Flash is even cheaper but may not match V4 Flash's quality for complex summarization.
Compare 95 AI Models Side by Side
DeepSeek V4 Flash is one of 95 models tracked on APIpulse. Compare pricing, context windows, and features across 11 providers.
Frequently Asked Questions
Related Pages
- DeepSeek V4 Pro Pricing — Output-optimized at $0.66/$1.98
- DeepSeek V4 Overview — Both Flash and Pro compared
- Qwen 3.7 Flash Pricing — Cheapest model at $0.03/$0.13
- GPT-5.4 nano Pricing — OpenAI's budget option at $0.20/$1.25
- Full Model Rankings — All 95 models ranked by price