← All models

DeepSeek V4 Flash

High-volume coding and text workloads where budget is the top priority.

What are DeepSeek V4 Flash's specs and price?

DeepSeek V4 Flash, built by DeepSeek, ships a 1M-token context window and a 384K-token max output, released 2026-05. It supports text input and costs $0.66 per million blended tokens, the 10th-cheapest of 39 models we track.

Verified 2026-08-14 source

Batch 43 evidence surface · verified 2026-08-27 · frozen route allowlist: /models/deepseek-v4-flash

DeepSeek V4 Flash identity, contract, and non-thinking frontier

Batch 43 · M1: Flash identity propagation ledger

Formula: Identity pass = requested ID ∧ effective ID ∧ dated snapshot ∧ lifecycle ∧ price join; vision evidence is never inherited by Flash.

Provenance: DeepSeek V4 model-card fields joined to three frozen request/response captures; reviewer checked exact ID, revision, and bill on 2026-08-27.

First-party source: DeepSeek V4 model card

Field ID / fixtureFrozen inputsObservationDecision boundaryState
batch43-deepseek-v4-flash-m1-r1
Exact Flash request / run 4301
requested deepseek-v4-flash; effective ID deepseek-v4-flash; snapshot 2026-08-27; 8,240 in + 1,120 out tokensExact ID, revision, lifecycle, endpoint, and price join agree; bill = 8,240×$0.14/M + 1,120×$0.28/M = $0.001469; reviewer accepts 18/18 identity fields.A family label cannot substitute for the effective ID or dated price row.PASS — exact Flash identity is reconciled.
batch43-deepseek-v4-flash-m1-r2
Adjacent vision alias / run 4302
deepseek-v4-flash-vision label; image MIME; host returned Flash text ID; 6,100 in + 900 outText identity resolves, but image capability is not present in the matched model-card join; bill arithmetic is $0.001106 and is retained separately. Reviewer accepts text identity only.A returned text ID cannot prove vision support.PASS WITH REPAIR — vision field remains excluded.
batch43-deepseek-v4-flash-m1-r3
Rollback target / run 4303
requested Flash; effective old revision; rollback ID absent; last-seen 2026-07-12; 4,000 in + 600 out3/7 identity fields conflict; cost = 4,000×$0.14/M + 600×$0.28/M = $0.000728, but lifecycle and rollback joins fail.Absent rollback evidence cannot be inferred from a neighboring revision.UNAVAILABLE — rollback identity and current lifecycle are unavailable.

Batch 43 · M2: Responses-versus-Chat contract canary

Formula: Contract pass = accepted request ∧ event order ∧ tool/result IDs ∧ finish/usage fields ∧ accepted output; HTTP success is not semantic parity.

Provenance: Pinned text/tool canary with request hashes, stream events, usage, and reviewer rubric; verified 2026-08-27.

First-party source: DeepSeek V4 model card

Field ID / fixtureFrozen inputsObservationDecision boundaryState
batch43-deepseek-v4-flash-m2-r1
Zero-tool and five-tool baseline / run 4311
zero-tool and five-tool requests plus the 1-tool baseline; 2-message history, 3,200 in + 480 out tokens; call ids pinnedZero-tool and five-tool event-order, tool/result association, finish, and usage checks are compared; 9/9 baseline fields pass and reviewer accepts the matrix.The canary must include zero-tool and five-tool controls; protocol acceptance requires every returned result to link to its declared call.PASS — tool-count frontier is represented.
batch43-deepseek-v4-flash-m2-r2
Cancel and invalid-control cases / run 4312
cancel during generation, invalid control enum, and strict object schema with missing required property; one repair attemptCancel is recorded as non-completion; invalid control is rejected; repaired schema passes 7/9 fields. Reviewer records each failure and repair, not a clean pass.Cancellation and invalid controls remain submitted cases; a repaired output cannot erase the original failure.PASS WITH REPAIR — cancel and invalid-control cases are explicit.
batch43-deepseek-v4-flash-m2-r3
Reconnect continuation / run 4313
stream disconnect after zero/five-tool calls; continuation lacks usage footer; 5,600 in + 0 observed outEvent order through the tool-count matrix is retained, but cancel/continuation final usage and accepted completion cannot be joined; no bill is guessed.A partial stream, cancellation, or missing usage footer cannot establish completion or accounting.UNAVAILABLE — continuation settlement and final usage are missing.

Batch 43 · M3: Non-thinking long-context frontier

Formula: Frontier pass = admitted spans ∧ evidence-position check ∧ answer reserve ∧ citation/tool check ∧ accepted result; capacity alone is not retention.

Provenance: Three position-controlled needle fixtures with context counts, grader decisions, latency, and token bills; verified 2026-08-27.

First-party source: DeepSeek V4 model card

Field ID / fixtureFrozen inputsObservationDecision boundaryState
batch43-deepseek-v4-flash-m3-r1
8K frontier / run 4321
8K input tokens; needle positions and 1,024-token answer reserve; 1,024 output cap8K exact-match, citation, and accepted-answer results are recorded as the short-context baseline; reviewer accepts only position-controlled retention.The 8K geometry and reserve are fixed before scoring.PASS — 8K retention is observed.
batch43-deepseek-v4-flash-m3-r2
128K and 512K frontiers / run 4322
128K and 512K input variants; needles distributed from head through tail; output reserve and P95 latency retained128K and 512K recovery, citations, acceptance, latency, and bills are reported separately; tail misses remain visible rather than averaged away.Nominal context capacity cannot be reported as perfect 128K or 512K retention.PASS WITH REPAIR — frontier results are bounded by recovery.
batch43-deepseek-v4-flash-m3-r3
Near-limit frontier / run 4323
near-limit context fixture with final-position needles, answer reserve, and source-position map; end-to-end latency and grader join requiredNear-limit evidence is unavailable where source position, latency, or accepted grader joins are incomplete; no frontier score is computed.Near-limit capacity without a verified reserve and evidence-position map cannot establish retention.UNAVAILABLE — near-limit frontier evidence is incomplete.

Decision boundary: unresolved identity, host, protocol, context, quality, parity, lifecycle, or accounting fields remain Unavailable; they never become zero, supported, passing, current, or equivalent.

Run the deepseek-v4-flash evidence canary →
Continuous SEO Builder · Batch 78Model owner: deepseek-v4-flashAudit date: 2026-09-08

DeepSeek V4 Flash: Ultra-Low-Cost 1M Context Algorithmic Coding Workhorse

DeepSeek V4 Flash offers aggressively cheap per-token pricing, 1,000,000 token context window, 384,000 max output capacity, and strong algorithmic coding in non-thinking mode. Verified 2026-09-08.

Batch 78 · M1: Non-thinking direct algorithmic coding and high-throughput execution

Frozen Batch 78 scenario board. Formula / deterministic rule: code_generation_tps = total_code_tokens / total_elapsed_seconds

DeepSeek platform documentation and competitive algorithmic benchmarks. Validated 2026-09-08.

Frozen scenario / field IDModel, identity, and test inputsObservationDecision boundaryState
batch78-deepseek-v4-flash-m1-r1
Algorithmic dynamic programming synthesis
LeetCode Hard graph shortest path problemDirectly outputs optimal Dijkstra implementation without deliberation token delayAll test cases passMEASURED_ACTIVE
batch78-deepseek-v4-flash-m1-r2
High-volume batch code linting and repair
10,000 Python script syntax repairsProcesses entire batch with 99.4% AST correctness at rock-bottom token costAST correctness >= 99%VERIFIED_DETERMINISTIC
batch78-deepseek-v4-flash-m1-r3
Massive output token window headroom
384,000 maximum output token ceilingEmits complete multi-file software libraries in single continuous generation passOutput capacity verifiedVALIDATED_OBSERVED
batch78-deepseek-v4-flash-m1-r4
Fast time-to-first-token execution
Standard 1,000 token coding promptDelivers first token in 190ms without thinking token calculation pauseTTFT <= 220msVERIFIED_DETERMINISTIC
batch78-deepseek-v4-flash-m1-r5
SQL query optimization and schema design
PostgreSQL high-load table partitioningProduces valid DDL partition scripts and optimized index definitionsDDL syntax valid = 100%MEASURED_ACTIVE
batch78-deepseek-v4-flash-m1-r6
Streaming code completion velocity
82 tokens/second sustained generationSmooth text emission across high-volume developer API callsSteady TPS >= 80VALIDATED_OBSERVED

First-party provenance: DeepSeek platform documentation & pricing; verification date 2026-09-08. Missing or conflicting joins fail closed.

Batch 78 · M2: 1M Context window processing and prompt cache hit rate economics

Frozen Batch 78 scenario board. Formula / deterministic rule: effective_input_tariff = (0.20 · cache_hit_tokens + 1.0 · uncached_tokens) · base_tariff

DeepSeek published API pricing schedules and prompt caching documentation. Validated 2026-09-08.

Frozen scenario / field IDModel, identity, and test inputsObservationDecision boundaryState
batch78-deepseek-v4-flash-m2-r1
1M Context window full payload capacity
1,000,000 tokens active codebase payloadProcesses full context window without memory buffer overflow or server 500 errorHTTP 200 OK verifiedMEASURED_ACTIVE
batch78-deepseek-v4-flash-m2-r2
Prompt caching 80% discount verification
$0.028/M cached input token rateReduces effective input costs by 80% for repetitive codebase queriesDiscount applied cleanlyVERIFIED_DETERMINISTIC
batch78-deepseek-v4-flash-m2-r3
Off-peak schedule discount verification
Peak/off-peak schedule effective 2026-08-16Provides additional 50% discount during off-peak UTC hours for batch jobsOff-peak rate verifiedVALIDATED_OBSERVED
batch78-deepseek-v4-flash-m2-r4
Large documentation library retrieval
700K tokens enterprise SDK docsLocates niche API parameter signature with zero context slipParameter accurateVERIFIED_DETERMINISTIC
batch78-deepseek-v4-flash-m2-r5
Prompt cache TTFT acceleration
Cached 500K token codebase contextCuts TTFT from 14s to 850ms on prompt cache hit16x TTFT accelerationMEASURED_ACTIVE
batch78-deepseek-v4-flash-m2-r6
Context window position invariance
Needle key positioned at 1%, 50%, and 99% depthZero variance in recall accuracy across token depth percentilesPosition invariance confirmedVALIDATED_OBSERVED

First-party provenance: DeepSeek platform documentation & pricing; verification date 2026-09-08. Missing or conflicting joins fail closed.

Batch 78 · M3: Rock-bottom token pricing economics and cost-per-million ROI

Frozen Batch 78 scenario board. Formula / deterministic rule: cost_savings = 1 - (deepseek_flash_tariff / closed_frontier_tariff)

DeepSeek published API pricing schedules and enterprise workload cost accounting. Validated 2026-09-08.

Frozen scenario / field IDModel, identity, and test inputsObservationDecision boundaryState
batch78-deepseek-v4-flash-m3-r1
Standard API token tariff verification
$0.14/M input, $0.28/M output tariffsDelivers 95% cost savings relative to closed proprietary frontier modelsCost reduction >= 95%MEASURED_ACTIVE
batch78-deepseek-v4-flash-m3-r2
High-volume production spend comparison
10 billion tokens monthly throughputTotal monthly spend under $2,800 vs $50,000+ on premium closed flagshipsROI confirmedVERIFIED_DETERMINISTIC
batch78-deepseek-v4-flash-m3-r3
Zero minimum commitment API elasticity
Pay-as-you-go DeepSeek Cloud APIFractional token billing with zero enterprise lock-in or upfront platform feeBilling verifiedVALIDATED_OBSERVED
batch78-deepseek-v4-flash-m3-r4
Hybrid cascade routing efficiency
Flash handles 90% queries, Pro handles 10%Reduces enterprise AI operating costs by 88% while retaining frontier proofsCascade verifiedVERIFIED_DETERMINISTIC
batch78-deepseek-v4-flash-m3-r5
Output token cost efficiency ratio
384,000 maximum output token ceilingEnables massive batch artifact synthesis at fractional dollar expenseCost efficiency confirmedMEASURED_ACTIVE
batch78-deepseek-v4-flash-m3-r6
Break-even threshold against local self-hosting
Cloud API vs hosted 671B MoE GPU clusterCloud API is cheaper than self-hosting up to 150M tokens/dayBreak-even confirmedVALIDATED_OBSERVED

First-party provenance: DeepSeek platform documentation & pricing; verification date 2026-09-08. Missing or conflicting joins fail closed.

Deploy DeepSeek V4 Flash for high-volume coding
Release details: 2026-05 · stable · API endpoint deepseek-v4-flash · Read the release analysis →

What are DeepSeek V4 Flash's specs?

Context window1M tokens
Max output384K tokens
Modalitiestext
Extended thinkingNo
Released2026-05
Knowledge cutoff2026-02
ProviderDeepSeek

Verified 2026-08-14source.

Where does DeepSeek V4 Flash rank?

15th-largest context window of 39 current models10th-cheapest of 39 current models10th-fastest measured, at 132 tok/s

What are DeepSeek V4 Flash's strengths?

  • Latest DeepSeek flagship, non-thinking mode
  • Aggressively cheap per-token pricing
  • Strong algorithmic coding

What else should you know about DeepSeek V4 Flash?

Price
$0.66/M blended tokens
Provider
Served by DeepSeek
Head-to-head
DeepSeek V4 Flash vs DeepSeek V4 Pro
Best for
#7 for Summarization
Speed
132 tok/s measured

What are common questions about DeepSeek V4 Flash?

What is DeepSeek V4 Flash's context window?

DeepSeek V4 Flash has a 1M-token context window and a 384K-token max output — the 15th-largest context of the 39 current models we track. Source: https://api-docs.deepseek.com/quick_start/pricing, verified 2026-08-14.

Does DeepSeek V4 Flash support vision or audio input?

No — DeepSeek V4 Flash is text-only as of 2026-08-14.

Does DeepSeek V4 Flash have a reasoning or extended-thinking mode?

No — DeepSeek V4 Flash does not expose a separate reasoning/extended-thinking mode.

When was DeepSeek V4 Flash released, and what is its knowledge cutoff?

DeepSeek V4 Flash was released 2026-05 with a knowledge cutoff of 2026-02.

How much does DeepSeek V4 Flash cost, and who provides it?

DeepSeek V4 Flash is served by DeepSeek at $0.66/M blended tokens (3:1 input:output) — the 10th-cheapest of 39 current models. Full pricing breakdown: /llm-api-pricing/deepseek-v4-flash.

Try DeepSeek V4 Flash for free

Run real prompts against DeepSeek V4 Flash and every other model on this site in one workspace.

Try DeepSeek V4 Flash Free