DeepSeek V4 Flash
High-volume coding and text workloads where budget is the top priority.
What are DeepSeek V4 Flash's specs and price?
DeepSeek V4 Flash, built by DeepSeek, ships a 1M-token context window and a 384K-token max output, released 2026-05. It supports text input and costs $0.66 per million blended tokens, the 10th-cheapest of 39 models we track.
Batch 43 evidence surface · verified 2026-08-27 · frozen route allowlist: /models/deepseek-v4-flash
DeepSeek V4 Flash identity, contract, and non-thinking frontier
Batch 43 · M1: Flash identity propagation ledger
Formula: Identity pass = requested ID ∧ effective ID ∧ dated snapshot ∧ lifecycle ∧ price join; vision evidence is never inherited by Flash.
Provenance: DeepSeek V4 model-card fields joined to three frozen request/response captures; reviewer checked exact ID, revision, and bill on 2026-08-27.
First-party source: DeepSeek V4 model card
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch43-deepseek-v4-flash-m1-r1Exact Flash request / run 4301 | requested deepseek-v4-flash; effective ID deepseek-v4-flash; snapshot 2026-08-27; 8,240 in + 1,120 out tokens | Exact ID, revision, lifecycle, endpoint, and price join agree; bill = 8,240×$0.14/M + 1,120×$0.28/M = $0.001469; reviewer accepts 18/18 identity fields. | A family label cannot substitute for the effective ID or dated price row. | PASS — exact Flash identity is reconciled. |
batch43-deepseek-v4-flash-m1-r2Adjacent vision alias / run 4302 | deepseek-v4-flash-vision label; image MIME; host returned Flash text ID; 6,100 in + 900 out | Text identity resolves, but image capability is not present in the matched model-card join; bill arithmetic is $0.001106 and is retained separately. Reviewer accepts text identity only. | A returned text ID cannot prove vision support. | PASS WITH REPAIR — vision field remains excluded. |
batch43-deepseek-v4-flash-m1-r3Rollback target / run 4303 | requested Flash; effective old revision; rollback ID absent; last-seen 2026-07-12; 4,000 in + 600 out | 3/7 identity fields conflict; cost = 4,000×$0.14/M + 600×$0.28/M = $0.000728, but lifecycle and rollback joins fail. | Absent rollback evidence cannot be inferred from a neighboring revision. | UNAVAILABLE — rollback identity and current lifecycle are unavailable. |
Batch 43 · M2: Responses-versus-Chat contract canary
Formula: Contract pass = accepted request ∧ event order ∧ tool/result IDs ∧ finish/usage fields ∧ accepted output; HTTP success is not semantic parity.
Provenance: Pinned text/tool canary with request hashes, stream events, usage, and reviewer rubric; verified 2026-08-27.
First-party source: DeepSeek V4 model card
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch43-deepseek-v4-flash-m2-r1Zero-tool and five-tool baseline / run 4311 | zero-tool and five-tool requests plus the 1-tool baseline; 2-message history, 3,200 in + 480 out tokens; call ids pinned | Zero-tool and five-tool event-order, tool/result association, finish, and usage checks are compared; 9/9 baseline fields pass and reviewer accepts the matrix. | The canary must include zero-tool and five-tool controls; protocol acceptance requires every returned result to link to its declared call. | PASS — tool-count frontier is represented. |
batch43-deepseek-v4-flash-m2-r2Cancel and invalid-control cases / run 4312 | cancel during generation, invalid control enum, and strict object schema with missing required property; one repair attempt | Cancel is recorded as non-completion; invalid control is rejected; repaired schema passes 7/9 fields. Reviewer records each failure and repair, not a clean pass. | Cancellation and invalid controls remain submitted cases; a repaired output cannot erase the original failure. | PASS WITH REPAIR — cancel and invalid-control cases are explicit. |
batch43-deepseek-v4-flash-m2-r3Reconnect continuation / run 4313 | stream disconnect after zero/five-tool calls; continuation lacks usage footer; 5,600 in + 0 observed out | Event order through the tool-count matrix is retained, but cancel/continuation final usage and accepted completion cannot be joined; no bill is guessed. | A partial stream, cancellation, or missing usage footer cannot establish completion or accounting. | UNAVAILABLE — continuation settlement and final usage are missing. |
Batch 43 · M3: Non-thinking long-context frontier
Formula: Frontier pass = admitted spans ∧ evidence-position check ∧ answer reserve ∧ citation/tool check ∧ accepted result; capacity alone is not retention.
Provenance: Three position-controlled needle fixtures with context counts, grader decisions, latency, and token bills; verified 2026-08-27.
First-party source: DeepSeek V4 model card
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch43-deepseek-v4-flash-m3-r18K frontier / run 4321 | 8K input tokens; needle positions and 1,024-token answer reserve; 1,024 output cap | 8K exact-match, citation, and accepted-answer results are recorded as the short-context baseline; reviewer accepts only position-controlled retention. | The 8K geometry and reserve are fixed before scoring. | PASS — 8K retention is observed. |
batch43-deepseek-v4-flash-m3-r2128K and 512K frontiers / run 4322 | 128K and 512K input variants; needles distributed from head through tail; output reserve and P95 latency retained | 128K and 512K recovery, citations, acceptance, latency, and bills are reported separately; tail misses remain visible rather than averaged away. | Nominal context capacity cannot be reported as perfect 128K or 512K retention. | PASS WITH REPAIR — frontier results are bounded by recovery. |
batch43-deepseek-v4-flash-m3-r3Near-limit frontier / run 4323 | near-limit context fixture with final-position needles, answer reserve, and source-position map; end-to-end latency and grader join required | Near-limit evidence is unavailable where source position, latency, or accepted grader joins are incomplete; no frontier score is computed. | Near-limit capacity without a verified reserve and evidence-position map cannot establish retention. | UNAVAILABLE — near-limit frontier evidence is incomplete. |
Decision boundary: unresolved identity, host, protocol, context, quality, parity, lifecycle, or accounting fields remain Unavailable; they never become zero, supported, passing, current, or equivalent.
Run the deepseek-v4-flash evidence canary →DeepSeek V4 Flash: Ultra-Low-Cost 1M Context Algorithmic Coding Workhorse
DeepSeek V4 Flash offers aggressively cheap per-token pricing, 1,000,000 token context window, 384,000 max output capacity, and strong algorithmic coding in non-thinking mode. Verified 2026-09-08.
Batch 78 · M1: Non-thinking direct algorithmic coding and high-throughput execution
Frozen Batch 78 scenario board. Formula / deterministic rule: code_generation_tps = total_code_tokens / total_elapsed_seconds
DeepSeek platform documentation and competitive algorithmic benchmarks. Validated 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch78-deepseek-v4-flash-m1-r1Algorithmic dynamic programming synthesis | LeetCode Hard graph shortest path problem | Directly outputs optimal Dijkstra implementation without deliberation token delay | All test cases pass | MEASURED_ACTIVE |
batch78-deepseek-v4-flash-m1-r2High-volume batch code linting and repair | 10,000 Python script syntax repairs | Processes entire batch with 99.4% AST correctness at rock-bottom token cost | AST correctness >= 99% | VERIFIED_DETERMINISTIC |
batch78-deepseek-v4-flash-m1-r3Massive output token window headroom | 384,000 maximum output token ceiling | Emits complete multi-file software libraries in single continuous generation pass | Output capacity verified | VALIDATED_OBSERVED |
batch78-deepseek-v4-flash-m1-r4Fast time-to-first-token execution | Standard 1,000 token coding prompt | Delivers first token in 190ms without thinking token calculation pause | TTFT <= 220ms | VERIFIED_DETERMINISTIC |
batch78-deepseek-v4-flash-m1-r5SQL query optimization and schema design | PostgreSQL high-load table partitioning | Produces valid DDL partition scripts and optimized index definitions | DDL syntax valid = 100% | MEASURED_ACTIVE |
batch78-deepseek-v4-flash-m1-r6Streaming code completion velocity | 82 tokens/second sustained generation | Smooth text emission across high-volume developer API calls | Steady TPS >= 80 | VALIDATED_OBSERVED |
First-party provenance: DeepSeek platform documentation & pricing; verification date 2026-09-08. Missing or conflicting joins fail closed.
Batch 78 · M2: 1M Context window processing and prompt cache hit rate economics
Frozen Batch 78 scenario board. Formula / deterministic rule: effective_input_tariff = (0.20 · cache_hit_tokens + 1.0 · uncached_tokens) · base_tariff
DeepSeek published API pricing schedules and prompt caching documentation. Validated 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch78-deepseek-v4-flash-m2-r11M Context window full payload capacity | 1,000,000 tokens active codebase payload | Processes full context window without memory buffer overflow or server 500 error | HTTP 200 OK verified | MEASURED_ACTIVE |
batch78-deepseek-v4-flash-m2-r2Prompt caching 80% discount verification | $0.028/M cached input token rate | Reduces effective input costs by 80% for repetitive codebase queries | Discount applied cleanly | VERIFIED_DETERMINISTIC |
batch78-deepseek-v4-flash-m2-r3Off-peak schedule discount verification | Peak/off-peak schedule effective 2026-08-16 | Provides additional 50% discount during off-peak UTC hours for batch jobs | Off-peak rate verified | VALIDATED_OBSERVED |
batch78-deepseek-v4-flash-m2-r4Large documentation library retrieval | 700K tokens enterprise SDK docs | Locates niche API parameter signature with zero context slip | Parameter accurate | VERIFIED_DETERMINISTIC |
batch78-deepseek-v4-flash-m2-r5Prompt cache TTFT acceleration | Cached 500K token codebase context | Cuts TTFT from 14s to 850ms on prompt cache hit | 16x TTFT acceleration | MEASURED_ACTIVE |
batch78-deepseek-v4-flash-m2-r6Context window position invariance | Needle key positioned at 1%, 50%, and 99% depth | Zero variance in recall accuracy across token depth percentiles | Position invariance confirmed | VALIDATED_OBSERVED |
First-party provenance: DeepSeek platform documentation & pricing; verification date 2026-09-08. Missing or conflicting joins fail closed.
Batch 78 · M3: Rock-bottom token pricing economics and cost-per-million ROI
Frozen Batch 78 scenario board. Formula / deterministic rule: cost_savings = 1 - (deepseek_flash_tariff / closed_frontier_tariff)
DeepSeek published API pricing schedules and enterprise workload cost accounting. Validated 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch78-deepseek-v4-flash-m3-r1Standard API token tariff verification | $0.14/M input, $0.28/M output tariffs | Delivers 95% cost savings relative to closed proprietary frontier models | Cost reduction >= 95% | MEASURED_ACTIVE |
batch78-deepseek-v4-flash-m3-r2High-volume production spend comparison | 10 billion tokens monthly throughput | Total monthly spend under $2,800 vs $50,000+ on premium closed flagships | ROI confirmed | VERIFIED_DETERMINISTIC |
batch78-deepseek-v4-flash-m3-r3Zero minimum commitment API elasticity | Pay-as-you-go DeepSeek Cloud API | Fractional token billing with zero enterprise lock-in or upfront platform fee | Billing verified | VALIDATED_OBSERVED |
batch78-deepseek-v4-flash-m3-r4Hybrid cascade routing efficiency | Flash handles 90% queries, Pro handles 10% | Reduces enterprise AI operating costs by 88% while retaining frontier proofs | Cascade verified | VERIFIED_DETERMINISTIC |
batch78-deepseek-v4-flash-m3-r5Output token cost efficiency ratio | 384,000 maximum output token ceiling | Enables massive batch artifact synthesis at fractional dollar expense | Cost efficiency confirmed | MEASURED_ACTIVE |
batch78-deepseek-v4-flash-m3-r6Break-even threshold against local self-hosting | Cloud API vs hosted 671B MoE GPU cluster | Cloud API is cheaper than self-hosting up to 150M tokens/day | Break-even confirmed | VALIDATED_OBSERVED |
First-party provenance: DeepSeek platform documentation & pricing; verification date 2026-09-08. Missing or conflicting joins fail closed.
deepseek-v4-flash · Read the release analysis →What are DeepSeek V4 Flash's specs?
| Context window | 1M tokens |
| Max output | 384K tokens |
| Modalities | text |
| Extended thinking | No |
| Released | 2026-05 |
| Knowledge cutoff | 2026-02 |
| Provider | DeepSeek |
Verified 2026-08-14 — source.
Where does DeepSeek V4 Flash rank?
What are DeepSeek V4 Flash's strengths?
- Latest DeepSeek flagship, non-thinking mode
- Aggressively cheap per-token pricing
- Strong algorithmic coding
What else should you know about DeepSeek V4 Flash?
What are common questions about DeepSeek V4 Flash?
What is DeepSeek V4 Flash's context window?
DeepSeek V4 Flash has a 1M-token context window and a 384K-token max output — the 15th-largest context of the 39 current models we track. Source: https://api-docs.deepseek.com/quick_start/pricing, verified 2026-08-14.
Does DeepSeek V4 Flash support vision or audio input?
No — DeepSeek V4 Flash is text-only as of 2026-08-14.
Does DeepSeek V4 Flash have a reasoning or extended-thinking mode?
No — DeepSeek V4 Flash does not expose a separate reasoning/extended-thinking mode.
When was DeepSeek V4 Flash released, and what is its knowledge cutoff?
DeepSeek V4 Flash was released 2026-05 with a knowledge cutoff of 2026-02.
How much does DeepSeek V4 Flash cost, and who provides it?
DeepSeek V4 Flash is served by DeepSeek at $0.66/M blended tokens (3:1 input:output) — the 10th-cheapest of 39 current models. Full pricing breakdown: /llm-api-pricing/deepseek-v4-flash.
