DeepSeek V4 Pro API Pricing: 1M Context at $0.66/$1.98

DeepSeek V4 Pro offers 1M-token context with the lowest output price of any model in its class. At $1.98/M output, it's 3x cheaper than DeepSeek V4 Flash for generation-heavy workloads.

Updated Aug 13, 2026 · 120 models tracked across 16 providers

TL;DR

⚠️ Peak/Off-Peak Pricing (since Aug 16, 2026): DeepSeek now uses time-based pricing. Off-peak (most hours): $0.66/$1.98 per million tokens. Peak (01:00–04:00 and 06:00–10:00 UTC): $1.32/$3.96 — double the off-peak rate. Since Aug 23: weekends (Sat/Sun Beijing time) are all-day off-peak. Prices below show off-peak rates. Cached input tokens are $0.003625/M regardless of time.

DeepSeek V4 Pro Pricing Breakdown

Off-peak pricing at different monthly volumes (peak rates are 2× higher):

Monthly Volume Off-Peak Input Off-Peak Output Off-Peak Total (50/50)
1M tokens$0.66$1.98$1.32
10M tokens$6.60$19.80$13.20
100M tokens$66.00$198.00$132.00
1B tokens$660.00$1,980.00$1,320.00

At 100M tokens/month off-peak with a 50/50 input/output split, DeepSeek V4 Pro costs $132. Peak hours double these rates. Schedule heavy workloads during off-peak hours to minimize costs.

DeepSeek V4 Family: Pro vs Flash

Model Input $/M Output $/M Context Best For
DeepSeek V4 Flash $0.22 $0.66 1M Input-heavy tasks (classification, summarization)
DeepSeek V4 Pro $0.66 $1.98 1M Output-heavy tasks (generation, coding)

Decision guide: If your workload is input-heavy (classify, summarize, extract), choose V4 Flash at $0.22/$0.66. If it's output-heavy (generate, write, code), choose V4 Pro at $0.66/$1.98 — the output savings (3x cheaper) often outweigh the input premium.

DeepSeek V4 Pro vs Other 1M-Context Models

Model Input $/M Output $/M Context Provider
Qwen 3.7 Flash $0.03 $0.13 1M Alibaba
Gemini 2.5 Flash-Lite $0.10 $0.40 1M Google
DeepSeek V4 Flash $0.22 $0.66 1M DeepSeek
DeepSeek V4 Pro $0.66 $1.98 1M DeepSeek
Gemini 2.5 Flash $0.30 $2.50 1M Google
GPT-5.6 Luna $0.20 $1.20 1.05M OpenAI
Claude Sonnet 5 $2.00 $10.00 1M Anthropic

DeepSeek V4 Pro has the lowest output price ($1.98) among all 1M-context models. Qwen 3.7 Flash is cheaper overall but has lower quality; GPT-5.6 Luna is comparable in quality but 38% more expensive on output.

When to Use DeepSeek V4 Pro

✍️ Content Generation

Generate articles, reports, or documentation at the lowest output cost. At $1.98/M output, it's the cheapest way to produce long-form text with 1M context.

💻 Code Generation

Generate code with full codebase context. 1M tokens fits entire repositories; low output pricing makes it ideal for code-heavy workloads.

💬 Chatbots

Build conversational AI where responses are longer than queries. The low output price means cheaper conversations at scale.

📝 Translation

Translate long documents where output length roughly matches input. The balanced input/output pricing makes it cost-effective for translation.

🔄 Data Transformation

Convert data between formats (JSON to CSV, code to documentation, structured to natural language). Output-heavy workloads benefit from the low output price.

📚 Summarization + Elaboration

Summarize long documents then elaborate on key points. 1M context handles the input; low output pricing makes the elaboration affordable.

Real-World Cost: Generating 50,000 Articles/Month

Suppose you're building a content generation service that produces 50,000 articles per month, averaging 1,000 input tokens (prompts) and 2,000 output tokens (articles) per request:

ModelInput CostOutput CostMonthly Total
DeepSeek V4 Pro$21.75$87.00$108.75
DeepSeek V4 Flash$7.00$28.00$35.00
Gemini 2.5 Flash$15.00$250.00$265.00
GPT-5.6 Luna$10.00$120.00$130.00
Claude Sonnet 5$100.00$1,000.00$1,100.00

DeepSeek V4 Pro at $108.75/month is cheaper than GPT-5.6 Luna ($130) and far cheaper than Gemini 2.5 Flash ($265) or Sonnet 5 ($1,100) for this output-heavy workload. V4 Flash is even cheaper if quality is sufficient.

Compare 95 AI Models Side by Side

DeepSeek V4 Pro is one of 120 models tracked on APIpulse. Compare pricing, context windows, and features across 16 providers.

Frequently Asked Questions

How much does DeepSeek V4 Pro cost?
DeepSeek V4 Pro costs $0.66 per million input tokens and $1.98 per million output tokens. This gives it the lowest output price of any model with 1M context — ideal for output-heavy workloads.
What is the context window of DeepSeek V4 Pro?
DeepSeek V4 Pro has a 1M token context window. This matches the largest context windows available from Google, OpenAI, and Anthropic, letting you process very long documents or codebases in a single request.
How does DeepSeek V4 Pro compare to DeepSeek V4 Flash?
V4 Pro ($0.66/$1.98) is 3x more expensive on input but 3x cheaper on output than V4 Flash ($0.22/$0.66). V4 Pro is better for output-heavy tasks; V4 Flash is better for input-heavy tasks like classification or summarization.
Is DeepSeek V4 Pro cheaper than GPT-5.4 nano?
On input, GPT-5.4 nano ($0.20) is cheaper than V4 Pro ($0.66). On output, V4 Pro ($1.98) is cheaper than GPT-5.4 nano ($1.25). V4 Pro also has 1M context vs 400K for nano. For output-heavy tasks with long context, V4 Pro is the better value.
Is DeepSeek V4 Pro available on all platforms?
DeepSeek V4 Pro is available directly through DeepSeek's API and through third-party providers like Together.ai, Fireworks, and others. It's not available on major cloud platforms like AWS Bedrock or Google Vertex AI.

Related Pages