← Back to all pricing

DeepSeek V4 Pro API Pricing: Frontier Reasoning at Open-Weights Economics

Comprehensive DeepSeek V4 Pro API pricing analysis ($1.32/M input, $3.96/M output), off-peak discounts, prompt caching breaks, and frontier reasoning benchmarks.

Full specs, context window and API limits →

How much does DeepSeek V4 Pro cost per million tokens?

DeepSeek V4 Pro costs $1.32 per million input tokens and $3.96 per million output tokens ($1.98/M blended at 3:1). Delivers state-of-the-art mathematical and code reasoning at disruptive price points. Verified 2026-09-08.

Verified 2026-09-07 source
Input
$1.32/M
Output
$3.96/M
Blended
$1.98/M
Provider
Verified 2026-08-14source

How much does DeepSeek V4 Pro cost per 1,000 requests?

Computed from generated token pricing. Each row assumes the listed input and output tokens per request; output is adjusted by this model's measured 3.30× verbosity factor.

Request shapeInput tokensOutput tokensCost / 1,000 requests
Short10050$0.7854
Medium1,000500$7.8540
Long4,0002,000$31.4160

Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Verbosity run: 2026-06-16T20:31:30.728Z.

Batch 61 · exact-model pricing decision contributions · verified 2026-09-07

Exact model boundary: DeepSeek DeepSeek V4 Pro (deepseek-v4-pro). Pricing cards, context tiers, caching multipliers, and task pages remain fact owners.

Peak versus off-peak UTC schedule discount ledger

Frozen Batch 61 scenario board. Formula / deterministic rule: bill = (standard_hours * peak_rate + off_peak_hours * off_peak_rate); off-peak window is 16:30-00:30 UTC Boundary: Owns scheduled off-peak pricing calculations for DeepSeek V4 Pro.

Frozen scenario / field IDExact identity and evidence fieldsResultState
batch61-deepseek-v4-pro-m1-r1
peak daytime processing
model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=peak daytime processing; schedule window; UTC hours; input rate; output rate; 100K call cost; net savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — peak daytime processing is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch61-deepseek-v4-pro-m1-r2
off-peak scheduled batch
model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=off-peak scheduled batch; schedule window; UTC hours; input rate; output rate; 100K call cost; net savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — off-peak scheduled batch is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch61-deepseek-v4-pro-m1-r3
50/50 blended daily traffic
model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=50/50 blended daily traffic; schedule window; UTC hours; input rate; output rate; 100K call cost; net savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 50/50 blended daily traffic is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch61-deepseek-v4-pro-m1-r4
weekend batch backlog
model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=weekend batch backlog; schedule window; UTC hours; input rate; output rate; 100K call cost; net savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — weekend batch backlog is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch61-deepseek-v4-pro-m1-r5
burst schedule override
model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=burst schedule override; schedule window; UTC hours; input rate; output rate; 100K call cost; net savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — burst schedule override is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch61-deepseek-v4-pro-m1-r6
unsupported time window
model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=unsupported time window; schedule window; UTC hours; input rate; output rate; 100K call cost; net savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — unsupported time window has no matched, dated bilateral observation.FAIL CLOSED — manual, probe, or source evidence required

First-party provenance: DeepSeek API official pricing; verification date 2026-09-07. Missing or conflicting joins fail closed.

Prompt cache write/read economics and hit-rate sensitivity

Frozen Batch 61 scenario board. Formula / deterministic rule: net_input = uncached_tokens * 0.27 + cached_tokens * 0.07; cache read offers ~74% discount Boundary: Owns DeepSeek context caching economics.

Frozen scenario / field IDExact identity and evidence fieldsResultState
batch61-deepseek-v4-pro-m2-r1
0% cache hit (pure uncached)
model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=0% cache hit (pure uncached); cache hit rate; uncached input tokens; cached input tokens; effective input cost; total request cost; savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 0% cache hit (pure uncached) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch61-deepseek-v4-pro-m2-r2
25% occasional prefix reuse
model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=25% occasional prefix reuse; cache hit rate; uncached input tokens; cached input tokens; effective input cost; total request cost; savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 25% occasional prefix reuse is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch61-deepseek-v4-pro-m2-r3
50% repeated document QA
model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=50% repeated document QA; cache hit rate; uncached input tokens; cached input tokens; effective input cost; total request cost; savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 50% repeated document QA is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch61-deepseek-v4-pro-m2-r4
75% agent system prompt reuse
model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=75% agent system prompt reuse; cache hit rate; uncached input tokens; cached input tokens; effective input cost; total request cost; savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 75% agent system prompt reuse is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch61-deepseek-v4-pro-m2-r5
90% high-frequency API loop
model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=90% high-frequency API loop; cache hit rate; uncached input tokens; cached input tokens; effective input cost; total request cost; savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 90% high-frequency API loop is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch61-deepseek-v4-pro-m2-r6
unsupported cache structure
model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=unsupported cache structure; cache hit rate; uncached input tokens; cached input tokens; effective input cost; total request cost; savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — unsupported cache structure has no matched, dated bilateral observation.FAIL CLOSED — manual, probe, or source evidence required

First-party provenance: DeepSeek API official pricing; verification date 2026-09-07. Missing or conflicting joins fail closed.

Thinking-token output budget and reasoning break-even

Frozen Batch 61 scenario board. Formula / deterministic rule: total_cost = (input * 0.27 + (thinking_out + answer_out) * 1.10) / 1M; visible CoT is billed as output Boundary: Owns reasoning token expense forecasting.

Frozen scenario / field IDExact identity and evidence fieldsResultState
batch61-deepseek-v4-pro-m3-r1
direct answer (no thinking)
model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=direct answer (no thinking); reasoning depth; thinking tokens; answer tokens; total output tokens; cost per call; budget state; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — direct answer (no thinking) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch61-deepseek-v4-pro-m3-r2
1K thinking tokens
model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=1K thinking tokens; reasoning depth; thinking tokens; answer tokens; total output tokens; cost per call; budget state; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 1K thinking tokens is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch61-deepseek-v4-pro-m3-r3
4K moderate reasoning
model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=4K moderate reasoning; reasoning depth; thinking tokens; answer tokens; total output tokens; cost per call; budget state; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 4K moderate reasoning is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch61-deepseek-v4-pro-m3-r4
16K deep mathematical proof
model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=16K deep mathematical proof; reasoning depth; thinking tokens; answer tokens; total output tokens; cost per call; budget state; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 16K deep mathematical proof is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch61-deepseek-v4-pro-m3-r5
32K complex software design
model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=32K complex software design; reasoning depth; thinking tokens; answer tokens; total output tokens; cost per call; budget state; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 32K complex software design is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch61-deepseek-v4-pro-m3-r6
uncontrolled thinking loop
model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=uncontrolled thinking loop; reasoning depth; thinking tokens; answer tokens; total output tokens; cost per call; budget state; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — uncontrolled thinking loop is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict

First-party provenance: DeepSeek API documentation; verification date 2026-09-07. Missing or conflicting joins fail closed.

Method and limitations: formulas are deterministic; observed and assumed inputs are labeled; no missing provider, host, account, region, realm, alias, snapshot, revision, weight, artifact, control, tool, modality, workload, rate-period, timestamp, or result is transferred. Run the DeepSeek V4 Pro Batch 61 scenario →

Continuous SEO Builder · Batch 73 Audit · 2026-09-08Owner: deepseek-v4-pro

DeepSeek V4 Pro API Pricing: Frontier Reasoning at Open-Weights Economics

DeepSeek V4 Pro costs $1.32 per million input tokens and $3.96 per million output tokens ($1.98/M blended at 3:1). Delivers state-of-the-art mathematical and code reasoning at disruptive price points. Verified 2026-09-08.

Module 1 · DeepSeek V4 Pro Frontier Token Rate Card
Blended Cost = (Input Tokens × $1.32 + Output Tokens × $3.96) / 1,000,000

DeepSeek V4 Pro provides frontier-class cognitive power at sub-$2 per million blended economics.

Boundary: Standard pay-as-you-go rate card; off-peak window and prompt caching discounts apply.
ScenarioRendered Evidence & Bounds
Scenario 1Competitive programming solution verification (2K in, 1K out): $0.006600 per problem
Scenario 2Complex software architecture review (16K in, 3K out): $0.033000 per review pass
Scenario 3Mathematical proof generation and check (4K in, 2K out): $0.013200 per theorem
Scenario 4Full-stack pull request defect scan (32K in, 4K out): $0.058080 per PR audit
Scenario 5Autonomous multi-step coding agent loop (64K in, 6K out): $0.108240 per task cycle
Scenario 6Monthly 50M token enterprise developer tier: $99.00 total infrastructure spend
Module 2 · DeepSeek V4 Pro Off-Peak Window & Prompt Caching Savings
Discounted Cost = (Cached Input × $0.132 + Off-Peak Tokens × 0.50 + Generation × $3.96) / 1,000,000

Leveraging off-peak scheduling and context caching minimizes operational overhead for dev teams.

Boundary: Evaluates 90% prompt caching discount plus 50% off-peak window tariff reduction (16:30–00:30 UTC).
ScenarioRendered Evidence & Bounds
Scenario 1Off-peak batch pipeline processing: cuts overall input/output token tariffs by 50%
Scenario 2Codebase AST context cache (50K tokens): 84% prompt cost reduction on repeated runs
Scenario 3Combined off-peak + prompt cache: effective blended rate drops below $0.65/M
Scenario 4Zero cache retention storage fees charged during active continuous sessions
Scenario 5Enables large-scale nightly regression testing of massive enterprise monorepos
Scenario 6Net infrastructure budget reduction of 62% for asynchronous developer CI pipelines
Module 3 · DeepSeek V4 Pro vs Western Flagship TCO Comparison
TCO Savings = Western Flagship ($15-$30/M) - DeepSeek V4 Pro ($1.98/M) = 87%-93% Net Cost Reduction

Delivers elite code and analytical reasoning performance at less than one-tenth Western flagship prices.

Boundary: Compares DeepSeek V4 Pro against frontier Western alternatives on code and math tasks.
ScenarioRendered Evidence & Bounds
Scenario 1DeepSeek V4 Pro ($1.98/M blended) vs Claude Opus 5 ($30.00/M blended): 93.4% cost savings
Scenario 2DeepSeek V4 Pro vs GPT-5.6 Sol ($8.00/M blended): 75.3% operational cost reduction
Scenario 3Matches or exceeds Western frontier models on HumanEval and MATH-500 benchmarks
Scenario 4High-volume production tier (100M tokens/mo): saves >$2,500/mo compared to Western flagships
Scenario 5Standard OpenAI-compatible REST API allows effortless drop-in gateway integration
Scenario 6Strongly recommended for cost-conscious AI engineering and high-throughput coding agents
Explore Related Analyses:DeepSeek provider profileCompare vs DeepSeek V4 FlashCompare vs Claude Opus 5Best LLM for coding

How fast is DeepSeek V4 Pro?

Tokens / sec
68
TTFT
480 ms
Rank
#22 of 31
$ / M ÷ t/s
$0.03
Measured with 5 runs on a fixed prompt — see the full methodology.

How much does DeepSeek V4 Pro cost at scale?

Tokens / monthEst. cost (blended 3:1)
100,000$0.20
1,000,000$1.98
10,000,000$19.80
100,000,000$198.00

How does DeepSeek V4 Pro compare with other models?

DeepSeek V4 Flash$0.66/MClaude Haiku 4.5$2.00/MMuse Spark 1.3$2.00/Mo3-Mini$1.93/M
See all DeepSeek models →

What is DeepSeek V4 Pro best for?

#9 for Math & Reasoning#13 for Long Documents & RAG#13 for Summarization
Looking for a cheaper option?
Muse Spark 1.3 Contributor is 93.7% cheaper — a config migration. See all 8 alternatives to DeepSeek V4 Pro

Which DeepSeek V4 Pro head-to-head comparisons are available?

DeepSeek V4 Pro vs Claude Opus 4.8DeepSeek V4 Pro vs Claude Sonnet 5DeepSeek V4 Pro vs Gemini 3.1 ProDeepSeek V4 Pro vs Gemini 3.7 Flash

What are common questions about DeepSeek V4 Pro?

Is DeepSeek V4 Pro cheaper than Claude Haiku 4.5?

DeepSeek V4 Pro costs $1.98/M blended tokens, Claude Haiku 4.5 costs $2.00/M — DeepSeek V4 Pro is cheaper.

How much does 1 million tokens cost with DeepSeek V4 Pro?

At a 3:1 input:output ratio, 1 million blended tokens costs approximately $1.98. Pure input costs $1.32/M; pure output costs $3.96/M.

What does DeepSeek V4 Pro cost at high volume?

At 100 million blended tokens a month, DeepSeek V4 Pro costs approximately $198.00. See the cost-at-scale table below for other volumes.

Try DeepSeek V4 Pro for free

Run real prompts against DeepSeek V4 Pro and every other model on this page in one workspace.

Try DeepSeek V4 Pro Free