DeepSeek V4 Pro API pricing (August 2026)
Base rate — bills 2x during Beijing peak hours (09:00-12:00 and 14:00-18:00 UTC+8).
What it costs at real volumes
A chat-shaped workload: 800 input tokens and 300 output tokens per request, no caching, no batch.
Specifications
- Provider
- DeepSeek
- Context window
- 1.0M tokens
- Max output
- 384K tokens
- Tier
- frontier
- Status
- GA
- Blended $/1M (3:1)
- $0.544
- Cache break-even hit rate
- 0% — caching always saves
- Prices verified
- 4 Aug 2026
Source: DeepSeek pricing. Finitizer is not affiliated with DeepSeek.
Similar models
Other deepseek models
- DeepSeek V4 Flash$-0.369
Closest prices elsewhere
- Llama 3.3 70B (Groq)+$0.096
- GPT-5 mini+$0.144
- GPT-4.1 mini+$0.156
- Mistral Large 3+$0.206
Frequently asked questions
How much does DeepSeek V4 Pro cost per 1M tokens?
DeepSeek V4 Pro costs $0.435 per million input tokens and $0.870 per million output tokens at DeepSeek list prices. Output is 2.0x the input rate, so the input:output ratio of your workload drives the real cost more than the headline price.
Does DeepSeek V4 Pro support prompt caching?
Yes. Cached input is billed at $0.00363 per million tokens, 120x cheaper than uncached input. There is no cache-write premium, so caching is never a net loss.
What is the batch discount for DeepSeek V4 Pro?
DeepSeek V4 Pro does not offer a batch pricing tier.
What is DeepSeek V4 Pro's context window?
1.0M tokens of context with up to 384K output tokens per request. A larger context window raises the ceiling on what you can send, not the price per token — but filling it on every request is one of the most common causes of an unexpected bill.
Running DeepSeek V4 Pro in production?
Finitizer TokenOps attributes real token spend to teams, features, and prompts — so you find out which workload moved the bill before the invoice does.