← Back to all pricing

DeepSeek V4 Flash API Pricing: Ultra-Low-Cost High-Speed Intelligence

Comprehensive DeepSeek V4 Flash API pricing analysis ($0.44/M input, $1.32/M output), off-peak discounts, classification speed, and prompt caching breaks.

Full specs, context window and API limits →

How much does DeepSeek V4 Flash cost per million tokens?

DeepSeek V4 Flash costs $0.44 per million input tokens and $1.32 per million output tokens ($0.66/M blended at 3:1). A high-efficiency open-weights model designed for rapid classification, document summarization, and large-scale data structuring. Verified 2026-09-08.

Verified 2026-09-07 source
Input
$0.44/M
Output
$1.32/M
Blended
$0.66/M
Provider
Verified 2026-08-14source

How much does DeepSeek V4 Flash cost per 1,000 requests?

Computed from generated token pricing. Each row assumes the listed input and output tokens per request; output is adjusted by this model's measured 2.59× verbosity factor.

Request shapeInput tokensOutput tokensCost / 1,000 requests
Short10050$0.2149
Medium1,000500$2.1494
Long4,0002,000$8.5976

Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Verbosity run: 2026-06-16T20:31:30.728Z.

Three model-specific pricing decisions

Flash is non-thinking mode, so the schedule and accepted-answer boundary stay separate from provider-wide scheduling and the broad Flash/Pro comparison.

1. UTC peak/off-peak plus cache-hit monthly schedule

UTC trafficCache statusFormulaDecision
Peak allocation pHit/miss separatep × peak + (1 − p) × off-peakMeasure p in UTC
Off-peak allocation 1 − pHit/miss separate100K calls × applicable input/output ratesSchedule flexible work
Dated rateModel price recordCurrent Flash rate onlyProvider schedule source required for numeric discount

2. Flash versus Pro cost-per-accepted-answer boundary

ModeToken cost / 100KAccepted-answer rateBoundary
DeepSeek V4 Flash$237.60UnavailableThinking premium cannot be inferred
DeepSeek V4 Pro$712.80UnavailableRecord accepted answers before switching

3. Operational sensitivity

VariableDated valueKeep separate from token price
Retry rateUnavailableMultiply full request cost when measured
LatencyUnavailableDo not convert milliseconds to dollars
SLA / availabilityUnavailableNo SLA claim in pricing record

Verified 2026-08-14. Luna is the data owner for this rendered decision module. “Unavailable” means the current dated registry has no model-specific evidence; it is not a zero. First-party price source · Run this scenario in the playground.

All results are server-rendered for DeepSeek V4 Flash; formulas expose fixed inputs and missing evidence remains visibly unavailable.

Batch 62 · exact-model pricing decision contributions · verified 2026-09-07

Exact model boundary: DeepSeek DeepSeek V4 Flash (deepseek-v4-flash). Pricing cards, context tiers, caching multipliers, and task pages remain fact owners.

Prompt cache hit versus cache miss cost ledger

Frozen Batch 62 scenario board. Formula / deterministic rule: cost = (cache_miss_in * 0.44 + cache_hit_in * 0.11 + out * 1.32) / 1M; cache hit saves 75% Boundary: Owns DeepSeek V4 Flash prompt caching economics.

Frozen scenario / field IDExact identity and evidence fieldsResultState
batch62-deepseek-v4-flash-m1-r1
100% cache miss cold prompt
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=100% cache miss cold prompt; prompt tokens; cache hit percentage; output tokens; net cost; effective input rate; savings percentage; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 100% cache miss cold prompt is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch62-deepseek-v4-flash-m1-r2
50% cache hit warm prompt
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=50% cache hit warm prompt; prompt tokens; cache hit percentage; output tokens; net cost; effective input rate; savings percentage; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 50% cache hit warm prompt is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch62-deepseek-v4-flash-m1-r3
80% cache hit enterprise system prompt
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=80% cache hit enterprise system prompt; prompt tokens; cache hit percentage; output tokens; net cost; effective input rate; savings percentage; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 80% cache hit enterprise system prompt is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch62-deepseek-v4-flash-m1-r4
95% cache hit document Q&A loop
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=95% cache hit document Q&A loop; prompt tokens; cache hit percentage; output tokens; net cost; effective input rate; savings percentage; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 95% cache hit document Q&A loop is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch62-deepseek-v4-flash-m1-r5
cache eviction on cold start
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=cache eviction on cold start; prompt tokens; cache hit percentage; output tokens; net cost; effective input rate; savings percentage; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — cache eviction on cold start is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch62-deepseek-v4-flash-m1-r6
unsupported multi-part schema
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=unsupported multi-part schema; prompt tokens; cache hit percentage; output tokens; net cost; effective input rate; savings percentage; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — unsupported multi-part schema has no matched, dated bilateral observation.FAIL CLOSED — manual, probe, or source evidence required

First-party provenance: DeepSeek API official pricing; verification date 2026-09-07. Missing or conflicting joins fail closed.

High-volume classification and ETL pipeline budgeting

Frozen Batch 62 scenario board. Formula / deterministic rule: pipeline_cost = records * ((doc_tokens * 0.44 + json_out * 1.32) / 1M) Boundary: Owns high-volume data pipeline economics for DeepSeek V4 Flash.

Frozen scenario / field IDExact identity and evidence fieldsResultState
batch62-deepseek-v4-flash-m2-r1
100K structured web scrape records
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=100K structured web scrape records; record count; avg input tokens; avg output tokens; monthly cost; cost per 1K records; pipeline status; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 100K structured web scrape records is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch62-deepseek-v4-flash-m2-r2
500K customer sentiment reviews
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=500K customer sentiment reviews; record count; avg input tokens; avg output tokens; monthly cost; cost per 1K records; pipeline status; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 500K customer sentiment reviews is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch62-deepseek-v4-flash-m2-r3
1M log classification events
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=1M log classification events; record count; avg input tokens; avg output tokens; monthly cost; cost per 1K records; pipeline status; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 1M log classification events is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch62-deepseek-v4-flash-m2-r4
10M enterprise data enrichment batch
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=10M enterprise data enrichment batch; record count; avg input tokens; avg output tokens; monthly cost; cost per 1K records; pipeline status; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 10M enterprise data enrichment batch is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch62-deepseek-v4-flash-m2-r5
rate limit throttling backup
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=rate limit throttling backup; record count; avg input tokens; avg output tokens; monthly cost; cost per 1K records; pipeline status; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — rate limit throttling backup is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch62-deepseek-v4-flash-m2-r6
unparseable JSON output retry
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=unparseable JSON output retry; record count; avg input tokens; avg output tokens; monthly cost; cost per 1K records; pipeline status; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — unparseable JSON output retry is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict

First-party provenance: DeepSeek API documentation; verification date 2026-09-07. Missing or conflicting joins fail closed.

DeepSeek V4 Flash vs Pro thinking mode economic trade-off

Frozen Batch 62 scenario board. Formula / deterministic rule: cost_ratio = pro_cost / flash_cost; evaluates when reasoning tokens justify cost Boundary: Owns Flash vs Pro model routing decision boundaries.

Frozen scenario / field IDExact identity and evidence fieldsResultState
batch62-deepseek-v4-flash-m3-r1
simple classification (Flash optimal)
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=simple classification (Flash optimal); workload type; Flash cost; Pro cost; cost multiplier; accuracy delta threshold; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — simple classification (Flash optimal) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch62-deepseek-v4-flash-m3-r2
data formatting & translation (Flash optimal)
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=data formatting & translation (Flash optimal); workload type; Flash cost; Pro cost; cost multiplier; accuracy delta threshold; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — data formatting & translation (Flash optimal) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch62-deepseek-v4-flash-m3-r3
multi-step logical reasoning (Pro optimal)
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=multi-step logical reasoning (Pro optimal); workload type; Flash cost; Pro cost; cost multiplier; accuracy delta threshold; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — multi-step logical reasoning (Pro optimal) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch62-deepseek-v4-flash-m3-r4
complex math verification (Pro optimal)
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=complex math verification (Pro optimal); workload type; Flash cost; Pro cost; cost multiplier; accuracy delta threshold; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — complex math verification (Pro optimal) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch62-deepseek-v4-flash-m3-r5
hybrid router: 80% Flash / 20% Pro
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=hybrid router: 80% Flash / 20% Pro; workload type; Flash cost; Pro cost; cost multiplier; accuracy delta threshold; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — hybrid router: 80% Flash / 20% Pro is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch62-deepseek-v4-flash-m3-r6
untested reasoning requirement
model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=untested reasoning requirement; workload type; Flash cost; Pro cost; cost multiplier; accuracy delta threshold; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — untested reasoning requirement is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict

First-party provenance: DeepSeek API official pricing; verification date 2026-09-07. Missing or conflicting joins fail closed.

Method and limitations: formulas are deterministic; observed and assumed inputs are labeled; no missing provider, host, account, region, realm, alias, snapshot, revision, weight, artifact, control, tool, modality, workload, rate-period, timestamp, or result is transferred. Run the DeepSeek V4 Flash Batch 62 scenario →

Continuous SEO Builder · Batch 73 Audit · 2026-09-08Owner: deepseek-v4-flash

DeepSeek V4 Flash API Pricing: Ultra-Low-Cost High-Speed Intelligence

DeepSeek V4 Flash costs $0.44 per million input tokens and $1.32 per million output tokens ($0.66/M blended at 3:1). A high-efficiency open-weights model designed for rapid classification, document summarization, and large-scale data structuring. Verified 2026-09-08.

Module 1 · DeepSeek V4 Flash Micro-Cost Token Economics
Blended Cost = (Input Tokens × $0.44 + Output Tokens × $1.32) / 1,000,000

DeepSeek V4 Flash delivers high-throughput utility inference at $0.66/M blended tokens.

Boundary: Standard pay-as-you-go rate card; off-peak window and prompt caching discounts apply.
ScenarioRendered Evidence & Bounds
Scenario 1Customer ticket intent classification (800 in, 50 out): $0.000418 per ticket
Scenario 2E-commerce product specification extraction (1.5K in, 200 out): $0.000924 per product
Scenario 3Customer feedback sentiment scoring (2K in, 100 out): $0.001012 per review
Scenario 4Document summary paragraph generation (4K in, 300 out): $0.002156 per document
Scenario 5High-volume webhook data normalization (1K in, 100 out): $0.000572 per webhook
Scenario 6Monthly 100M token classification fleet: $66.00 total infrastructure spend
Module 2 · DeepSeek V4 Flash Off-Peak Tariff & Cache Amortization
Discounted Cost = (Cached Input × $0.044 + Off-Peak Tokens × 0.50 + Generation × $1.32) / 1,000,000

Combining off-peak execution with prompt caching delivers industry-leading data enrichment economics.

Boundary: Evaluates 90% prompt caching discount and 50% off-peak tariff reduction during off-peak hours.
ScenarioRendered Evidence & Bounds
Scenario 1Off-peak asynchronous data tagging: cuts baseline token prices by exactly 50%
Scenario 2Shared JSON extraction schema cache (8K tokens): 82% input cost reduction
Scenario 3Combined off-peak + prompt cache: effective blended rate drops below $0.25/M
Scenario 4Zero cache retention storage fees charged during active continuous sessions
Scenario 5Enables massive bulk data enrichment across multi-million record databases
Scenario 6Reduces enterprise data structuring operational costs by over 70%
Module 3 · DeepSeek V4 Flash High-Volume Batch Extraction Pipeline
Batch Efficiency = Throughput (tokens/sec) / Blended Price ($/M)

Exceptional token throughput pairs with rock-bottom pricing for high-volume enterprise ETL.

Boundary: Evaluates throughput optimization and concurrency scaling for automated enterprise workflows.
ScenarioRendered Evidence & Bounds
Scenario 1Processes 10,000 customer survey responses for under $5.00 total API cost
Scenario 2High streaming throughput (120+ tps) ensures zero queue delays on API gateways
Scenario 3Strict JSON mode adherence guarantees zero downstream serialization pipeline errors
Scenario 4Low-memory footprint supports massive concurrent connection limits on shared infrastructure
Scenario 5Reliable instruction following on complex multi-field schema extraction tasks
Scenario 6Ideal operational choice for high-volume ETL pipelines and real-time content moderation
Explore Related Analyses:DeepSeek provider profileCompare vs DeepSeek V4 ProCompare vs GPT-5.4 NanoCheapest AI API comparison

How fast is DeepSeek V4 Flash?

Tokens / sec
132
TTFT
280 ms
Rank
#10 of 31
$ / M ÷ t/s
$0.0050
Measured with 5 runs on a fixed prompt — see the full methodology.

How much does DeepSeek V4 Flash cost at scale?

Tokens / monthEst. cost (blended 3:1)
100,000$0.07
1,000,000$0.66
10,000,000$6.60
100,000,000$66.00

How does DeepSeek V4 Flash compare with other models?

DeepSeek V4 Pro$1.98/MGPT-5 Mini$0.69/MMistral Large 3$0.75/MGemini 3.1 Flash Lite$0.56/M
See all DeepSeek models →

What is DeepSeek V4 Flash best for?

#7 for Summarization#9 for Writing & Content#10 for Long Documents & RAG
Looking for a cheaper option?
Muse Spark 1.3 Contributor is 81.1% cheaper — a config migration. See all 8 alternatives to DeepSeek V4 Flash

Which DeepSeek V4 Flash head-to-head comparisons are available?

DeepSeek V4 Flash vs DeepSeek V4 Pro

What are common questions about DeepSeek V4 Flash?

Is DeepSeek V4 Flash cheaper than GPT-5 Mini?

DeepSeek V4 Flash costs $0.66/M blended tokens, GPT-5 Mini costs $0.69/M — DeepSeek V4 Flash is cheaper.

How much does 1 million tokens cost with DeepSeek V4 Flash?

At a 3:1 input:output ratio, 1 million blended tokens costs approximately $0.66. Pure input costs $0.44/M; pure output costs $1.32/M.

What does DeepSeek V4 Flash cost at high volume?

At 100 million blended tokens a month, DeepSeek V4 Flash costs approximately $66.00. See the cost-at-scale table below for other volumes.

Try DeepSeek V4 Flash for free

Run real prompts against DeepSeek V4 Flash and every other model on this page in one workspace.

Try DeepSeek V4 Flash Free