DeepSeek V4 Flash API Pricing: Ultra-Low-Cost High-Speed Intelligence
Comprehensive DeepSeek V4 Flash API pricing analysis ($0.44/M input, $1.32/M output), off-peak discounts, classification speed, and prompt caching breaks.
Full specs, context window and API limits →How much does DeepSeek V4 Flash cost per million tokens?
DeepSeek V4 Flash costs $0.44 per million input tokens and $1.32 per million output tokens ($0.66/M blended at 3:1). A high-efficiency open-weights model designed for rapid classification, document summarization, and large-scale data structuring. Verified 2026-09-08.
How much does DeepSeek V4 Flash cost per 1,000 requests?
Computed from generated token pricing. Each row assumes the listed input and output tokens per request; output is adjusted by this model's measured 2.59× verbosity factor.
| Request shape | Input tokens | Output tokens | Cost / 1,000 requests |
|---|---|---|---|
| Short | 100 | 50 | $0.2149 |
| Medium | 1,000 | 500 | $2.1494 |
| Long | 4,000 | 2,000 | $8.5976 |
Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Verbosity run: 2026-06-16T20:31:30.728Z.
Three model-specific pricing decisions
Flash is non-thinking mode, so the schedule and accepted-answer boundary stay separate from provider-wide scheduling and the broad Flash/Pro comparison.
1. UTC peak/off-peak plus cache-hit monthly schedule
| UTC traffic | Cache status | Formula | Decision |
|---|---|---|---|
| Peak allocation p | Hit/miss separate | p × peak + (1 − p) × off-peak | Measure p in UTC |
| Off-peak allocation 1 − p | Hit/miss separate | 100K calls × applicable input/output rates | Schedule flexible work |
| Dated rate | Model price record | Current Flash rate only | Provider schedule source required for numeric discount |
2. Flash versus Pro cost-per-accepted-answer boundary
| Mode | Token cost / 100K | Accepted-answer rate | Boundary |
|---|---|---|---|
| DeepSeek V4 Flash | $237.60 | Unavailable | Thinking premium cannot be inferred |
| DeepSeek V4 Pro | $712.80 | Unavailable | Record accepted answers before switching |
3. Operational sensitivity
| Variable | Dated value | Keep separate from token price |
|---|---|---|
| Retry rate | Unavailable | Multiply full request cost when measured |
| Latency | Unavailable | Do not convert milliseconds to dollars |
| SLA / availability | Unavailable | No SLA claim in pricing record |
Verified 2026-08-14. Luna is the data owner for this rendered decision module. “Unavailable” means the current dated registry has no model-specific evidence; it is not a zero. First-party price source · Run this scenario in the playground.
All results are server-rendered for DeepSeek V4 Flash; formulas expose fixed inputs and missing evidence remains visibly unavailable.
Batch 62 · exact-model pricing decision contributions · verified 2026-09-07
Exact model boundary: DeepSeek DeepSeek V4 Flash (deepseek-v4-flash). Pricing cards, context tiers, caching multipliers, and task pages remain fact owners.
Prompt cache hit versus cache miss cost ledger
Frozen Batch 62 scenario board. Formula / deterministic rule: cost = (cache_miss_in * 0.44 + cache_hit_in * 0.11 + out * 1.32) / 1M; cache hit saves 75% Boundary: Owns DeepSeek V4 Flash prompt caching economics.
| Frozen scenario / field ID | Exact identity and evidence fields | Result | State |
|---|---|---|---|
batch62-deepseek-v4-flash-m1-r1100% cache miss cold prompt | model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=100% cache miss cold prompt; prompt tokens; cache hit percentage; output tokens; net cost; effective input rate; savings percentage; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 100% cache miss cold prompt is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch62-deepseek-v4-flash-m1-r250% cache hit warm prompt | model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=50% cache hit warm prompt; prompt tokens; cache hit percentage; output tokens; net cost; effective input rate; savings percentage; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 50% cache hit warm prompt is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch62-deepseek-v4-flash-m1-r380% cache hit enterprise system prompt | model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=80% cache hit enterprise system prompt; prompt tokens; cache hit percentage; output tokens; net cost; effective input rate; savings percentage; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 80% cache hit enterprise system prompt is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch62-deepseek-v4-flash-m1-r495% cache hit document Q&A loop | model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=95% cache hit document Q&A loop; prompt tokens; cache hit percentage; output tokens; net cost; effective input rate; savings percentage; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 95% cache hit document Q&A loop is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch62-deepseek-v4-flash-m1-r5cache eviction on cold start | model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=cache eviction on cold start; prompt tokens; cache hit percentage; output tokens; net cost; effective input rate; savings percentage; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — cache eviction on cold start is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch62-deepseek-v4-flash-m1-r6unsupported multi-part schema | model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=unsupported multi-part schema; prompt tokens; cache hit percentage; output tokens; net cost; effective input rate; savings percentage; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — unsupported multi-part schema has no matched, dated bilateral observation. | FAIL CLOSED — manual, probe, or source evidence required |
First-party provenance: DeepSeek API official pricing; verification date 2026-09-07. Missing or conflicting joins fail closed.
High-volume classification and ETL pipeline budgeting
Frozen Batch 62 scenario board. Formula / deterministic rule: pipeline_cost = records * ((doc_tokens * 0.44 + json_out * 1.32) / 1M) Boundary: Owns high-volume data pipeline economics for DeepSeek V4 Flash.
| Frozen scenario / field ID | Exact identity and evidence fields | Result | State |
|---|---|---|---|
batch62-deepseek-v4-flash-m2-r1100K structured web scrape records | model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=100K structured web scrape records; record count; avg input tokens; avg output tokens; monthly cost; cost per 1K records; pipeline status; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 100K structured web scrape records is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch62-deepseek-v4-flash-m2-r2500K customer sentiment reviews | model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=500K customer sentiment reviews; record count; avg input tokens; avg output tokens; monthly cost; cost per 1K records; pipeline status; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 500K customer sentiment reviews is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch62-deepseek-v4-flash-m2-r31M log classification events | model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=1M log classification events; record count; avg input tokens; avg output tokens; monthly cost; cost per 1K records; pipeline status; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 1M log classification events is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch62-deepseek-v4-flash-m2-r410M enterprise data enrichment batch | model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=10M enterprise data enrichment batch; record count; avg input tokens; avg output tokens; monthly cost; cost per 1K records; pipeline status; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 10M enterprise data enrichment batch is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch62-deepseek-v4-flash-m2-r5rate limit throttling backup | model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=rate limit throttling backup; record count; avg input tokens; avg output tokens; monthly cost; cost per 1K records; pipeline status; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — rate limit throttling backup is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch62-deepseek-v4-flash-m2-r6unparseable JSON output retry | model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=unparseable JSON output retry; record count; avg input tokens; avg output tokens; monthly cost; cost per 1K records; pipeline status; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — unparseable JSON output retry is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
First-party provenance: DeepSeek API documentation; verification date 2026-09-07. Missing or conflicting joins fail closed.
DeepSeek V4 Flash vs Pro thinking mode economic trade-off
Frozen Batch 62 scenario board. Formula / deterministic rule: cost_ratio = pro_cost / flash_cost; evaluates when reasoning tokens justify cost Boundary: Owns Flash vs Pro model routing decision boundaries.
| Frozen scenario / field ID | Exact identity and evidence fields | Result | State |
|---|---|---|---|
batch62-deepseek-v4-flash-m3-r1simple classification (Flash optimal) | model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=simple classification (Flash optimal); workload type; Flash cost; Pro cost; cost multiplier; accuracy delta threshold; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — simple classification (Flash optimal) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch62-deepseek-v4-flash-m3-r2data formatting & translation (Flash optimal) | model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=data formatting & translation (Flash optimal); workload type; Flash cost; Pro cost; cost multiplier; accuracy delta threshold; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — data formatting & translation (Flash optimal) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch62-deepseek-v4-flash-m3-r3multi-step logical reasoning (Pro optimal) | model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=multi-step logical reasoning (Pro optimal); workload type; Flash cost; Pro cost; cost multiplier; accuracy delta threshold; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — multi-step logical reasoning (Pro optimal) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch62-deepseek-v4-flash-m3-r4complex math verification (Pro optimal) | model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=complex math verification (Pro optimal); workload type; Flash cost; Pro cost; cost multiplier; accuracy delta threshold; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — complex math verification (Pro optimal) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch62-deepseek-v4-flash-m3-r5hybrid router: 80% Flash / 20% Pro | model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=hybrid router: 80% Flash / 20% Pro; workload type; Flash cost; Pro cost; cost multiplier; accuracy delta threshold; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — hybrid router: 80% Flash / 20% Pro is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch62-deepseek-v4-flash-m3-r6untested reasoning requirement | model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=untested reasoning requirement; workload type; Flash cost; Pro cost; cost multiplier; accuracy delta threshold; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — untested reasoning requirement is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
First-party provenance: DeepSeek API official pricing; verification date 2026-09-07. Missing or conflicting joins fail closed.
Method and limitations: formulas are deterministic; observed and assumed inputs are labeled; no missing provider, host, account, region, realm, alias, snapshot, revision, weight, artifact, control, tool, modality, workload, rate-period, timestamp, or result is transferred. Run the DeepSeek V4 Flash Batch 62 scenario →
deepseek-v4-flashDeepSeek V4 Flash API Pricing: Ultra-Low-Cost High-Speed Intelligence
DeepSeek V4 Flash costs $0.44 per million input tokens and $1.32 per million output tokens ($0.66/M blended at 3:1). A high-efficiency open-weights model designed for rapid classification, document summarization, and large-scale data structuring. Verified 2026-09-08.
DeepSeek V4 Flash delivers high-throughput utility inference at $0.66/M blended tokens.
| Scenario | Rendered Evidence & Bounds |
|---|---|
| Scenario 1 | Customer ticket intent classification (800 in, 50 out): $0.000418 per ticket |
| Scenario 2 | E-commerce product specification extraction (1.5K in, 200 out): $0.000924 per product |
| Scenario 3 | Customer feedback sentiment scoring (2K in, 100 out): $0.001012 per review |
| Scenario 4 | Document summary paragraph generation (4K in, 300 out): $0.002156 per document |
| Scenario 5 | High-volume webhook data normalization (1K in, 100 out): $0.000572 per webhook |
| Scenario 6 | Monthly 100M token classification fleet: $66.00 total infrastructure spend |
Combining off-peak execution with prompt caching delivers industry-leading data enrichment economics.
| Scenario | Rendered Evidence & Bounds |
|---|---|
| Scenario 1 | Off-peak asynchronous data tagging: cuts baseline token prices by exactly 50% |
| Scenario 2 | Shared JSON extraction schema cache (8K tokens): 82% input cost reduction |
| Scenario 3 | Combined off-peak + prompt cache: effective blended rate drops below $0.25/M |
| Scenario 4 | Zero cache retention storage fees charged during active continuous sessions |
| Scenario 5 | Enables massive bulk data enrichment across multi-million record databases |
| Scenario 6 | Reduces enterprise data structuring operational costs by over 70% |
Exceptional token throughput pairs with rock-bottom pricing for high-volume enterprise ETL.
| Scenario | Rendered Evidence & Bounds |
|---|---|
| Scenario 1 | Processes 10,000 customer survey responses for under $5.00 total API cost |
| Scenario 2 | High streaming throughput (120+ tps) ensures zero queue delays on API gateways |
| Scenario 3 | Strict JSON mode adherence guarantees zero downstream serialization pipeline errors |
| Scenario 4 | Low-memory footprint supports massive concurrent connection limits on shared infrastructure |
| Scenario 5 | Reliable instruction following on complex multi-field schema extraction tasks |
| Scenario 6 | Ideal operational choice for high-volume ETL pipelines and real-time content moderation |
How fast is DeepSeek V4 Flash?
How much does DeepSeek V4 Flash cost at scale?
| Tokens / month | Est. cost (blended 3:1) |
|---|---|
| 100,000 | $0.07 |
| 1,000,000 | $0.66 |
| 10,000,000 | $6.60 |
| 100,000,000 | $66.00 |
How does DeepSeek V4 Flash compare with other models?
What is DeepSeek V4 Flash best for?
What should you explore next for DeepSeek V4 Flash?
Which DeepSeek V4 Flash head-to-head comparisons are available?
What are common questions about DeepSeek V4 Flash?
Is DeepSeek V4 Flash cheaper than GPT-5 Mini?
DeepSeek V4 Flash costs $0.66/M blended tokens, GPT-5 Mini costs $0.69/M — DeepSeek V4 Flash is cheaper.
How much does 1 million tokens cost with DeepSeek V4 Flash?
At a 3:1 input:output ratio, 1 million blended tokens costs approximately $0.66. Pure input costs $0.44/M; pure output costs $1.32/M.
What does DeepSeek V4 Flash cost at high volume?
At 100 million blended tokens a month, DeepSeek V4 Flash costs approximately $66.00. See the cost-at-scale table below for other volumes.
