DeepSeek V4 Pro API Pricing: Frontier Reasoning at Open-Weights Economics
Comprehensive DeepSeek V4 Pro API pricing analysis ($1.32/M input, $3.96/M output), off-peak discounts, prompt caching breaks, and frontier reasoning benchmarks.
Full specs, context window and API limits →How much does DeepSeek V4 Pro cost per million tokens?
DeepSeek V4 Pro costs $1.32 per million input tokens and $3.96 per million output tokens ($1.98/M blended at 3:1). Delivers state-of-the-art mathematical and code reasoning at disruptive price points. Verified 2026-09-08.
How much does DeepSeek V4 Pro cost per 1,000 requests?
Computed from generated token pricing. Each row assumes the listed input and output tokens per request; output is adjusted by this model's measured 3.30× verbosity factor.
| Request shape | Input tokens | Output tokens | Cost / 1,000 requests |
|---|---|---|---|
| Short | 100 | 50 | $0.7854 |
| Medium | 1,000 | 500 | $7.8540 |
| Long | 4,000 | 2,000 | $31.4160 |
Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Verbosity run: 2026-06-16T20:31:30.728Z.
Batch 61 · exact-model pricing decision contributions · verified 2026-09-07
Exact model boundary: DeepSeek DeepSeek V4 Pro (deepseek-v4-pro). Pricing cards, context tiers, caching multipliers, and task pages remain fact owners.
Peak versus off-peak UTC schedule discount ledger
Frozen Batch 61 scenario board. Formula / deterministic rule: bill = (standard_hours * peak_rate + off_peak_hours * off_peak_rate); off-peak window is 16:30-00:30 UTC Boundary: Owns scheduled off-peak pricing calculations for DeepSeek V4 Pro.
| Frozen scenario / field ID | Exact identity and evidence fields | Result | State |
|---|---|---|---|
batch61-deepseek-v4-pro-m1-r1peak daytime processing | model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=peak daytime processing; schedule window; UTC hours; input rate; output rate; 100K call cost; net savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — peak daytime processing is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch61-deepseek-v4-pro-m1-r2off-peak scheduled batch | model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=off-peak scheduled batch; schedule window; UTC hours; input rate; output rate; 100K call cost; net savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — off-peak scheduled batch is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch61-deepseek-v4-pro-m1-r350/50 blended daily traffic | model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=50/50 blended daily traffic; schedule window; UTC hours; input rate; output rate; 100K call cost; net savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 50/50 blended daily traffic is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch61-deepseek-v4-pro-m1-r4weekend batch backlog | model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=weekend batch backlog; schedule window; UTC hours; input rate; output rate; 100K call cost; net savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — weekend batch backlog is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch61-deepseek-v4-pro-m1-r5burst schedule override | model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=burst schedule override; schedule window; UTC hours; input rate; output rate; 100K call cost; net savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — burst schedule override is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch61-deepseek-v4-pro-m1-r6unsupported time window | model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=unsupported time window; schedule window; UTC hours; input rate; output rate; 100K call cost; net savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — unsupported time window has no matched, dated bilateral observation. | FAIL CLOSED — manual, probe, or source evidence required |
First-party provenance: DeepSeek API official pricing; verification date 2026-09-07. Missing or conflicting joins fail closed.
Prompt cache write/read economics and hit-rate sensitivity
Frozen Batch 61 scenario board. Formula / deterministic rule: net_input = uncached_tokens * 0.27 + cached_tokens * 0.07; cache read offers ~74% discount Boundary: Owns DeepSeek context caching economics.
| Frozen scenario / field ID | Exact identity and evidence fields | Result | State |
|---|---|---|---|
batch61-deepseek-v4-pro-m2-r10% cache hit (pure uncached) | model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=0% cache hit (pure uncached); cache hit rate; uncached input tokens; cached input tokens; effective input cost; total request cost; savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 0% cache hit (pure uncached) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch61-deepseek-v4-pro-m2-r225% occasional prefix reuse | model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=25% occasional prefix reuse; cache hit rate; uncached input tokens; cached input tokens; effective input cost; total request cost; savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 25% occasional prefix reuse is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch61-deepseek-v4-pro-m2-r350% repeated document QA | model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=50% repeated document QA; cache hit rate; uncached input tokens; cached input tokens; effective input cost; total request cost; savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 50% repeated document QA is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch61-deepseek-v4-pro-m2-r475% agent system prompt reuse | model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=75% agent system prompt reuse; cache hit rate; uncached input tokens; cached input tokens; effective input cost; total request cost; savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 75% agent system prompt reuse is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch61-deepseek-v4-pro-m2-r590% high-frequency API loop | model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=90% high-frequency API loop; cache hit rate; uncached input tokens; cached input tokens; effective input cost; total request cost; savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 90% high-frequency API loop is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch61-deepseek-v4-pro-m2-r6unsupported cache structure | model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=unsupported cache structure; cache hit rate; uncached input tokens; cached input tokens; effective input cost; total request cost; savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — unsupported cache structure has no matched, dated bilateral observation. | FAIL CLOSED — manual, probe, or source evidence required |
First-party provenance: DeepSeek API official pricing; verification date 2026-09-07. Missing or conflicting joins fail closed.
Thinking-token output budget and reasoning break-even
Frozen Batch 61 scenario board. Formula / deterministic rule: total_cost = (input * 0.27 + (thinking_out + answer_out) * 1.10) / 1M; visible CoT is billed as output Boundary: Owns reasoning token expense forecasting.
| Frozen scenario / field ID | Exact identity and evidence fields | Result | State |
|---|---|---|---|
batch61-deepseek-v4-pro-m3-r1direct answer (no thinking) | model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=direct answer (no thinking); reasoning depth; thinking tokens; answer tokens; total output tokens; cost per call; budget state; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — direct answer (no thinking) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch61-deepseek-v4-pro-m3-r21K thinking tokens | model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=1K thinking tokens; reasoning depth; thinking tokens; answer tokens; total output tokens; cost per call; budget state; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 1K thinking tokens is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch61-deepseek-v4-pro-m3-r34K moderate reasoning | model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=4K moderate reasoning; reasoning depth; thinking tokens; answer tokens; total output tokens; cost per call; budget state; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 4K moderate reasoning is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch61-deepseek-v4-pro-m3-r416K deep mathematical proof | model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=16K deep mathematical proof; reasoning depth; thinking tokens; answer tokens; total output tokens; cost per call; budget state; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 16K deep mathematical proof is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch61-deepseek-v4-pro-m3-r532K complex software design | model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=32K complex software design; reasoning depth; thinking tokens; answer tokens; total output tokens; cost per call; budget state; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 32K complex software design is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch61-deepseek-v4-pro-m3-r6uncontrolled thinking loop | model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=uncontrolled thinking loop; reasoning depth; thinking tokens; answer tokens; total output tokens; cost per call; budget state; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — uncontrolled thinking loop is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
First-party provenance: DeepSeek API documentation; verification date 2026-09-07. Missing or conflicting joins fail closed.
Method and limitations: formulas are deterministic; observed and assumed inputs are labeled; no missing provider, host, account, region, realm, alias, snapshot, revision, weight, artifact, control, tool, modality, workload, rate-period, timestamp, or result is transferred. Run the DeepSeek V4 Pro Batch 61 scenario →
deepseek-v4-proDeepSeek V4 Pro API Pricing: Frontier Reasoning at Open-Weights Economics
DeepSeek V4 Pro costs $1.32 per million input tokens and $3.96 per million output tokens ($1.98/M blended at 3:1). Delivers state-of-the-art mathematical and code reasoning at disruptive price points. Verified 2026-09-08.
DeepSeek V4 Pro provides frontier-class cognitive power at sub-$2 per million blended economics.
| Scenario | Rendered Evidence & Bounds |
|---|---|
| Scenario 1 | Competitive programming solution verification (2K in, 1K out): $0.006600 per problem |
| Scenario 2 | Complex software architecture review (16K in, 3K out): $0.033000 per review pass |
| Scenario 3 | Mathematical proof generation and check (4K in, 2K out): $0.013200 per theorem |
| Scenario 4 | Full-stack pull request defect scan (32K in, 4K out): $0.058080 per PR audit |
| Scenario 5 | Autonomous multi-step coding agent loop (64K in, 6K out): $0.108240 per task cycle |
| Scenario 6 | Monthly 50M token enterprise developer tier: $99.00 total infrastructure spend |
Leveraging off-peak scheduling and context caching minimizes operational overhead for dev teams.
| Scenario | Rendered Evidence & Bounds |
|---|---|
| Scenario 1 | Off-peak batch pipeline processing: cuts overall input/output token tariffs by 50% |
| Scenario 2 | Codebase AST context cache (50K tokens): 84% prompt cost reduction on repeated runs |
| Scenario 3 | Combined off-peak + prompt cache: effective blended rate drops below $0.65/M |
| Scenario 4 | Zero cache retention storage fees charged during active continuous sessions |
| Scenario 5 | Enables large-scale nightly regression testing of massive enterprise monorepos |
| Scenario 6 | Net infrastructure budget reduction of 62% for asynchronous developer CI pipelines |
Delivers elite code and analytical reasoning performance at less than one-tenth Western flagship prices.
| Scenario | Rendered Evidence & Bounds |
|---|---|
| Scenario 1 | DeepSeek V4 Pro ($1.98/M blended) vs Claude Opus 5 ($30.00/M blended): 93.4% cost savings |
| Scenario 2 | DeepSeek V4 Pro vs GPT-5.6 Sol ($8.00/M blended): 75.3% operational cost reduction |
| Scenario 3 | Matches or exceeds Western frontier models on HumanEval and MATH-500 benchmarks |
| Scenario 4 | High-volume production tier (100M tokens/mo): saves >$2,500/mo compared to Western flagships |
| Scenario 5 | Standard OpenAI-compatible REST API allows effortless drop-in gateway integration |
| Scenario 6 | Strongly recommended for cost-conscious AI engineering and high-throughput coding agents |
How fast is DeepSeek V4 Pro?
How much does DeepSeek V4 Pro cost at scale?
| Tokens / month | Est. cost (blended 3:1) |
|---|---|
| 100,000 | $0.20 |
| 1,000,000 | $1.98 |
| 10,000,000 | $19.80 |
| 100,000,000 | $198.00 |
How does DeepSeek V4 Pro compare with other models?
What is DeepSeek V4 Pro best for?
What should you explore next for DeepSeek V4 Pro?
Which DeepSeek V4 Pro head-to-head comparisons are available?
What are common questions about DeepSeek V4 Pro?
Is DeepSeek V4 Pro cheaper than Claude Haiku 4.5?
DeepSeek V4 Pro costs $1.98/M blended tokens, Claude Haiku 4.5 costs $2.00/M — DeepSeek V4 Pro is cheaper.
How much does 1 million tokens cost with DeepSeek V4 Pro?
At a 3:1 input:output ratio, 1 million blended tokens costs approximately $1.98. Pure input costs $1.32/M; pure output costs $3.96/M.
What does DeepSeek V4 Pro cost at high volume?
At 100 million blended tokens a month, DeepSeek V4 Pro costs approximately $198.00. See the cost-at-scale table below for other volumes.
