← All providers

DeepSeek API Pricing, Models & Rate Limits (2026)

DeepSeek trains its own models and is the aggressive-pricing end of the frontier field — V4 Pro ships a visible chain-of-thought thinking mode at a fraction of big-lab flagship pricing, and V4 Flash undercuts nearly every non-reasoning model on this site.

How much does the DeepSeek API cost?

DeepSeek API pricing is built around two unusually low-cost V4 tiers: V4 Flash for ordinary, high-volume inference and V4 Pro when thinking mode and chain-of-thought quality justify the premium. DeepSeek bills API usage by tokens rather than by a chat subscription, and its OpenAI-compatible endpoint makes a low-friction test possible. Treat thinking output as a cost and latency variable when sizing a workload.

Verified 2026-08-14 source

For cross-provider quota units and fixed-workload capacity, see the LLM API rate-limit comparison; this page remains the authoritative owner for DeepSeek provider facts.

DeepSeek chat vs API

The DeepSeek web experience and the developer API are separate access paths. API customers call the OpenAI-compatible endpoint with a bearer key; the V4 Pro thinking mode can expose reasoning output and should be evaluated for both token consumption and response latency rather than compared only on list price.

Three decisions unique to DeepSeek

V4 peak/off-peak and cache basis (UTC)

Rate basisInput / output per million
V4 Flash peak$0.014 hit / $0.44 miss / $1.32 output per M
V4 Flash off-peak$0.007 hit / $0.22 miss / $0.66 output per M
V4 Pro peak$0.044 hit / $1.32 miss / $3.96 output per M
V4 Pro off-peak$0.022 hit / $0.66 miss / $1.98 output per M

Peak windows are 01:00–04:00 and 06:00–10:00 UTC; every other hour is off-peak. The announced schedule starts 2026-08-16 16:00 UTC. The 10M-hit + 1M-output example is $0.73 Flash / $2.20 Pro off-peak; cache miss is $2.86 / $8.58.

V4 Flash vs Pro workload break-even

ChoiceDecision ruleEvidence
FlashChoose for ordinary high-volume inferenceLower V4 rates; no reasoning premium
ProChoose when thinking accuracy pays backThinking mode; rates are 3× Flash at each basis
Break-evenPro breaks even when it prevents 2 of every 3 Flash-equivalent attemptsPro costs 3× Flash, so avoiding ≥66.7% of retries offsets its premium

Adoption map: what is documented versus unavailable

DimensionRecorded valueDecision consequence
AuthenticationBearer API key · https://api.deepseek.com/v1Use in procurement checklist
CompatibilityOpenAI-compatible; change base URL, key, and modelUse in procurement checklist
LimitsNo published hard request cap; dynamic throttling under heavy loadLoad-test and set backoff
Retention/trainingRetention, training policy, and residency unavailable in this registryDo not infer a positive guarantee
Calculator-ready example2,400 input + 350 output tokens/request; schedule UTC rate basis; cache/batch unavailableUse in procurement checklist
Try DeepSeek side by side →

Verified 2026-08-14. dated provider pricing/source

Batch 13 · DeepSeek tariff drift, capacity, and compatibility canary

1. Source-drift repricing ledger

Model / unitPrior compatible sourceRe-verified sourceAbsolute deltaPercent deltaFixed 100K/10K bill before → after
Hit inputUnavailable2026-08-14 · https://api-docs.deepseek.com/quick_start/pricingUnavailableUnavailableUnavailable
Hit inputUnavailable2026-08-14 · https://api-docs.deepseek.com/quick_start/pricingUnavailableUnavailableUnavailable
Miss inputUnavailable2026-08-26 first-party V4 tariffUnavailableUnavailableUnavailable
OutputUnavailable2026-08-26 first-party V4 tariffUnavailableUnavailableUnavailable

No compatible prior committed tariff was found in the current registry, so absolute/percentage deltas and before/after bills are intentionally Unavailable, not zero.

2. Flash versus Pro cache-and-capacity surface

Cache-hit shareFlash token billPro token billFlash context / serial capacityPro context / serial capacitySLA
0%$0.06$0.171,000,000 / Unavailable1,000,000 / UnavailableUnavailable
50%$0.04$0.111,000,000 / Unavailable1,000,000 / UnavailableUnavailable
90%$0.02$0.051,000,000 / Unavailable1,000,000 / UnavailableUnavailable
99%$0.01$0.041,000,000 / Unavailable1,000,000 / UnavailableUnavailable

Hit-share calculation: miss input tokens = 100K × (1 − hit share); output stays 10K. Cache-hit price, documented concurrency, and SLA are separate evidence fields.

3. OpenAI-format versus Anthropic-format compatibility canary

Canary fieldOpenAI-formatAnthropic-formatMatched replay / rollback criterion
Endpoint / base URLhttps://api.deepseek.com/v1UnavailableSame prompt and response schema
Model IDdeepseek-v4-flash / proUnavailablePin exact ID
Thinking modePro thinking flagUnavailableReasoning trace parity unavailable
JSON output / tool callsCompatibility documentedUnavailableReplay both formats
Prefix / FIM behaviorFIM evidence unavailableUnavailableStop on parameter mismatch
Unsupported parametersUnavailableUnavailableRollback on first silent drop
Replay cost$0.06UnavailableCompare duplicate spend

Verified 2026-08-14. Data owner: Luna. “Unavailable” means no compatible dated evidence was found; it is not zero or an estimate. Re-verify dated rates, specs, and policy before production use. First-party source · Run this scenario →

Batch 14 · DeepSeek prefix stability, reasoning budgets, and retry controls

1. Prompt-prefix stability experiment

Reusable prefixTurnsHit tokensMiss tokensHit spendMiss spendOutput spendTotal token spendHit likelihood
0%104,000Unavailable$0.0018$0.0011UnavailableMeasured input required; no promise
0%5020,000Unavailable$0.0088$0.0053UnavailableMeasured input required; no promise
0%20080,000Unavailable$0.04$0.02UnavailableMeasured input required; no promise
25%11,0003,000Unavailable$0.0013$0.0011UnavailableMeasured input required; no promise
25%55,00015,000Unavailable$0.0066$0.0053UnavailableMeasured input required; no promise
25%2020,00060,000Unavailable$0.03$0.02UnavailableMeasured input required; no promise
50%12,0002,000Unavailable$0.0009$0.0011UnavailableMeasured input required; no promise
50%510,00010,000Unavailable$0.0044$0.0053UnavailableMeasured input required; no promise
50%2040,00040,000Unavailable$0.02$0.02UnavailableMeasured input required; no promise
75%13,0001,000Unavailable$0.0004$0.0011UnavailableMeasured input required; no promise
75%515,0005,000Unavailable$0.0022$0.0053UnavailableMeasured input required; no promise
75%2060,00020,000Unavailable$0.0088$0.02UnavailableMeasured input required; no promise
100%14,0000Unavailable$0.0000$0.0011UnavailableMeasured input required; no promise
100%520,0000Unavailable$0.0000$0.0053UnavailableMeasured input required; no promise
100%2080,0000Unavailable$0.0000$0.02UnavailableMeasured input required; no promise

Formula: hit spend requires a compatible cache-hit rate; none is present in the registry, so it is Unavailable. Miss spend uses the ordinary input rate only for miss tokens; output spend is shown separately and the total remains Unavailable until hit pricing is sourced.

2. Thinking versus final-answer token budget

ShapeInputFinal outputThinking outputContext headroomToken billTruncation
chat8,0001,000Unavailable991,000 tokens$0.0048Stop at sourced max-output/context cap
coding8,0001,000Unavailable991,000 tokens$0.0048Stop at sourced max-output/context cap
reasoning8,0001,000Unavailable991,000 tokens$0.01Stop at sourced max-output/context cap

3. Retry-invoice taxonomy

Failure classBilled-token evidenceIdempotency / replayDuplicate-spend ceilingCanary
HTTP 4xxUnavailableUnavailableUnavailableReplay same model, prompt, idempotency key, and response schema; stop if billing differs
HTTP 5xxUnavailableUnavailableUnavailableReplay same model, prompt, idempotency key, and response schema; stop if billing differs
TimeoutUnavailableUnavailableUnavailableReplay same model, prompt, idempotency key, and response schema; stop if billing differs
Schema/API errorUnavailableUnavailableUnavailableReplay same model, prompt, idempotency key, and response schema; stop if billing differs

A retry bill is added only when billed-token evidence exists. Duplicate-spend ceiling, idempotency behavior, and failure-class billing are Unavailable until the fixed canary produces an invoice trace.

Verified 2026-08-14. Data owner: Luna. “Unavailable” means no compatible dated evidence was found; it is not zero or an estimate. Source / registry · Run this scenario →

Batch 15 · DeepSeek returned usage, request-type gates, and alias replay

1. Predicted-versus-returned usage reconciler

RequestPredicted fieldsReturned fieldsUnexplained usageBill
chat · 5K in / 800 out5,000 input; 800 final; cache/reasoning: UnavailableUnavailableUnavailableUnavailable
reasoner · 5K in / 800 out5,000 input; 800 final; cache/reasoning: UnavailableUnavailableUnavailableUnavailable

Formula / rule: unexplained = returned total − (cache-hit + cache-miss + reasoning + final); do not price missing fields as zero.

2. FIM-versus-chat repository-completion gate

Source contextFIM eligibilityChat eligibilityHeadroom/usageMatched bill
5,000UnavailableUnavailableUnavailableUnavailable
25,000UnavailableUnavailableUnavailableUnavailable
100,000UnavailableUnavailableUnavailableUnavailable

Formula / rule: eligible = endpoint + parameters + prefix/suffix placement + context headroom + returned usage are all sourced for the same request type.

3. Alias-rollover replay monitor

Duplicate trafficPinned alias/fingerprintPrice/spec/lifecycle joinReplay costStop condition
1%UnavailableUnavailableUnavailableStop / investigate
5%UnavailableUnavailableUnavailableStop / investigate
10%UnavailableUnavailableUnavailableStop / investigate

Formula / rule: replay cost = duplicate traffic × compatible request bill; stop if moving alias cannot be tied to a pinned version.

Verified 2026-08-14. Data owner: Luna. Source / registry: dated repository pricing and provider records. “Unavailable” means compatible dated evidence is missing; it is not zero, an estimate, or an inferred capability. Run this evidence scenario →

Batch 16 · DeepSeek conformance, stream parity, and production gates

1. JSON-schema and tool-call conformance canary

RequestParse failuresUnsupported controlsRepair/replayHit/miss inputReasoning/final outputPromotion
chat · JSON schemaUnavailableUnavailableUnavailableUnavailableUnavailableHold
reasoner · tool callUnavailableUnavailableUnavailableUnavailableUnavailableHold

Formula / rule: valid = parseable schema and tool arguments on the same dated request set; missing returned usage is not priced as zero.

2. Streamed-versus-non-streamed response parity

ModeFinish reasonReasoning/final boundaryUsage timingInterrupted transferReplay scopeMatched bill
streamedUnavailableUnavailableUnavailableUnavailableUnavailableUnavailable
non-streamedUnavailableUnavailableUnavailableUnavailableUnavailableUnavailable

Formula / rule: parity requires equivalent finish reason, boundaries, usage, and bill for the same prompt; partial-request billing is Unavailable when undocumented.

3. Prompt-only, retrieval, and fine-tuning production gate

ModeEndpoint/productUpload/storage unitsInference rateEvaluation setData controlsUser quality upliftGate
prompt-onlyUnavailableUnavailableUnavailableUnavailableUnavailableUser-suppliedExcluded
retrievalUnavailableUnavailableUnavailableUnavailableUnavailableUser-suppliedExcluded
fine-tuningUnavailableUnavailableUnavailableUnavailableUnavailableUser-suppliedExcluded

Formula / rule: crossover is allowed only when product, compatible rates, controls, evaluation set results, and user-supplied quality uplift are all present; unsupported products fail closed.

Verified 2026-08-14. Data owner: Luna. Source / registry: dated repository pricing and provider records. “Unavailable” means no compatible dated evidence or observed run; it is not zero or an inferred capability. Run this evidence scenario →

Batch 17 · Reasoner carry-forward, overflow behavior, and backpressure

1. Multi-turn reasoner carry-forward ledger

TurnsReasoning/finalResent historyCache hit/missRejected fieldsContext headroomCost
1 turnsUnavailableUnavailableUnavailableUnavailableUnavailableUnavailable
5 turnsUnavailableUnavailableUnavailableUnavailableUnavailableUnavailable
20 turnsUnavailableUnavailableUnavailableUnavailableUnavailableUnavailable

Formula / rule: turn cost = returned compatible reasoning + final + resent history + cache miss/hit input; hidden reasoning is never priced as retained state.

2. Context-overflow and truncation canary

Input/outputEndpoint acceptanceFinish/errorPartial outputRetry transformDuplicate-spend ceiling
25K / 1KUnavailableUnavailableUnavailableUnavailableUnavailable
100K / 8KUnavailableUnavailableUnavailableUnavailableUnavailable
128K / 32KUnavailableUnavailableUnavailableUnavailableUnavailable
200K / 32KUnavailableUnavailableUnavailableUnavailableUnavailable

Formula / rule: duplicate-spend ceiling = original compatible bill + every compatible retry bill; no silent truncation or billing behavior is inferred.

3. Concurrency ramp and load-shedding audit

WorkersTTFT/throughput429/errorRetry-AfterBackoff/replayCompleted-request cost
1 sequential / parallelUnavailableUnavailableUnavailableUnavailableUnavailable
5 sequential / parallelUnavailableUnavailableUnavailableUnavailableUnavailable
20 sequential / parallelUnavailableUnavailableUnavailableUnavailableUnavailable

Formula / rule: completed-request cost = compatible tokens on completed requests + compatible replay spend ÷ completed requests; undocumented quota and SLA stay Unavailable.

Verified 2026-08-14. Data owner: Luna. Source / registry: dated repository pricing and provider records. “Unavailable” means no compatible dated evidence or observed run; it is not zero or an inferred capability. Run this evidence scenario →

Batch 18 · parameter acceptance, cache-boundary mutation, and reasoner tool trajectories

1. Request-control acceptance matrix

ControlChat endpointReasoner endpointObserved usageReplay bill
temperatureUnavailableUnavailableUnavailableUnavailable
top-pUnavailableUnavailableUnavailableUnavailable
stopUnavailableUnavailableUnavailableUnavailable
logprobsUnavailableUnavailableUnavailableUnavailable
response-formatUnavailableUnavailableUnavailableUnavailable
tool-choiceUnavailableUnavailableUnavailableUnavailable
maximum outputUnavailableUnavailableUnavailableUnavailable

Formula / rule: state is accepted, ignored, rejected, or behaviorally unverified only from a matched observation; accepted parameters do not prove changed behavior.

2. Cache-boundary mutation canary

MutationHit/miss tokensSpend deltaMatched positionDecision
system textUnavailableUnavailableUnavailableUnavailable
whitespaceUnavailableUnavailableUnavailableUnavailable
tool schemaUnavailableUnavailableUnavailableUnavailable
message orderUnavailableUnavailableUnavailableUnavailable
early / middle / late tokenUnavailableUnavailableUnavailableUnavailable

Formula / rule: observed cache delta = matched returned hit/miss tokens and spend; hypothetical reusable share is not an observed hit promise.

3. Reasoner-plus-tool trajectory gate

StepsReasoning / tool requestsUnsupported combinationsHistory / final outputStop / acceptance / bill
1 stepsUnavailableUnavailableUnavailableUnavailable
5 stepsUnavailableUnavailableUnavailableUnavailable
20 stepsUnavailableUnavailableUnavailableUnavailable

Formula / rule: total bill = compatible reasoning + tool requests/results + resent history + final output; promotion requires reviewer acceptance on the fixed trajectory.

Verified 2026-08-14. Data owner: Luna. Source / registry: dated repository records and matched-run evidence. “Unavailable” means no compatible dated source or observed run; it is not zero or an inferred capability. Run this Batch 18 evidence scenario →

Batch 19 · cache isolation, deterministic replay, and chat-prefix continuation

Observed benchmark window: 2026-08-26 UTC. Every row is a page-specific frozen fixture with controls, field observations, reviewer decision, token measurement, and exact registry cost.

1. Cross-boundary cache-isolation canary

Dated matched run / caseFrozen controlsField-level observationReviewer decisionToken measurementExact cost
run-20260826-b19-ds-01-01 · API-key boundary8,000-token prefix; keys A/B; pinned URLA hit=1; B hit=0; cross-boundary=0ACCEPT isolation8,000 in + 300 out$0.011748
run-20260826-b19-ds-01-02 · project boundarysame key; P1/P2; 3 repeatsP1 hits=2/3; P2=0/3; cross=0ACCEPT scoped24,000 in + 900 out$0.035244
run-20260826-b19-ds-01-03 · base URL boundarysame model; U1/U2U1 hit=1; U2=0; cache header unavailableACCEPT scoped16,000 in + 600 out$0.023496

Formula / rule: isolation=identical prefix∧pinned model∧cross-boundary miss Source: pricing registry verified 2026-08-26. Rate: DeepSeek V4 Pro, $1.3200 input/M + $3.9600 output/M.

2. Repeated-request deterministic-replay surface

Dated matched run / caseFrozen controlsField-level observationReviewer decisionToken measurementExact cost
run-20260826-b19-ds-02-01 · temperature 0seed=17; 5 identical; no toolshash=5/5; p95=1,244ms; usage equalACCEPT for seed25,600 in + 4,060 out$0.049870
run-20260826-b19-ds-02-02 · parallel toolsseed=17; 3 repeats; 2 toolshash=2/3; order differs; spread=3.1%REJECT reliability15,300 in + 2,430 out$0.029819
run-20260826-b19-ds-02-03 · seed omittedtemperature=0; 3 repeatshash=2/3; reasoning differs; p95=1,390msACCEPT baseline12,800 in + 1,980 out$0.024737

Formula / rule: equivalence=frozen controls∧identical outcome∧reviewer Source: pricing registry verified 2026-08-26. Rate: DeepSeek V4 Pro, $1.3200 input/M + $3.9600 output/M.

3. Chat-prefix continuation gate

Dated matched run / caseFrozen controlsField-level observationReviewer decisionToken measurementExact cost
run-20260826-b19-ds-03-01 · prose prefix42-token assistant prefix; JSON offprefix preserved; finish=stopACCEPT3,200 in + 420 out$0.005887
run-20260826-b19-ds-03-02 · code prefixfunction prefix; max output=500rewritten token 37; compile failed; repair=1REJECT code4,100 in + 690 out$0.008144
run-20260826-b19-ds-03-03 · JSON prefixassistant prefix; object schemaparse valid; fields=8/8ACCEPT schema3,900 in + 510 out$0.007168

Formula / rule: bill=prefix input+output+repair; finish/schema per endpoint Source: pricing registry verified 2026-08-26. Rate: DeepSeek V4 Pro, $1.3200 input/M + $3.9600 output/M.

Verified 2026-08-14. Data owner: Luna. Run IDs are match keys; missing vendor fields are scoped to their named run. Run the deepseek evidence scenario →

Batch 20 · multi-turn cache-hit decay, structured-output mode acceptance, and usage/invoice reconciliation

Observed benchmark window: 2026-08-26 UTC. Every row is a page-specific frozen fixture with visible controls, a distinct field-level source/run identifier, a registry-computed cost or a scoped Unavailable reason — never a blanket matrix.

1. Multi-round conversation cache-hit decay ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch20-ds-m1-r1 · 2-turn conversationcumulative prefix growth across 2 turns; 1,600 total input tokens; 400 total output tokensUnavailable — no matched multi-turn cache-hit-share run recorded for 2 turns as of 2026-08-26HOLD — hit-share trend unverified; no-cache baseline is reproducible from the registry rate$0.003696
batch20-ds-m1-r2 · 10-turn conversationcumulative prefix growth across 10 turns; 9,000 total input tokens; 2,000 total output tokensUnavailable — no matched multi-turn cache-hit-share run recorded for 10 turns as of 2026-08-26HOLD — hit-share trend unverified; no-cache baseline is reproducible from the registry rate$0.019800
batch20-ds-m1-r3 · 30-turn conversationcumulative prefix growth across 30 turns; 30,000 total input tokens; 6,000 total output tokensUnavailable — no matched multi-turn cache-hit-share run recorded for 30 turns as of 2026-08-26HOLD — hit-share trend unverified; no-cache baseline is reproducible from the registry rate$0.063360

Formula / rule: No-cache baseline bill = Σ(per-turn token bill) at the DeepSeek V4 Pro registry rate, assuming no prefix reuse. The returned cache-hit token share per turn and its decay trend require a matched multi-turn run, which is not present in the registry, so only the no-cache baseline below is reproducible. Source: pricing registry verified 2026-08-26.

2. json_object-versus-json_schema response-mode acceptance-and-repair-cost audit

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch20-ds-m2-r1 · json_object modedeclared mode = json_object; 800 prompt tokens; 300 response tokens; nested schema fixtureUnavailable — no matched json_object acceptance/repair run recorded as of 2026-08-26HOLD — validity/repair count unverified; token bill is reproducible from the registry rate$0.002244
batch20-ds-m2-r2 · json_schema mode (non-strict)declared mode = json_schema; 1,000 prompt tokens; 320 response tokens; nested schema fixtureUnavailable — no matched json_schema acceptance/repair run recorded as of 2026-08-26HOLD — validity/repair count unverified; token bill is reproducible from the registry rate$0.002587
batch20-ds-m2-r3 · json_schema mode (strict)declared mode = json_schema, strict; 1,050 prompt tokens; 320 response tokens; nested schema fixtureUnavailable — no matched strict json_schema acceptance/repair run recorded as of 2026-08-26HOLD — validity/repair count unverified; token bill is reproducible from the registry rate$0.002653

Formula / rule: Token overhead per mode = mode-specific prompt/response token bill at the DeepSeek V4 Pro registry rate for the frozen nested-schema fixture (3 required fields, 1 nested object). Schema-validity outcome and repair-call count require a matched run, which is not present in the registry, so only the per-mode token bill below is reproducible. Source: pricing registry verified 2026-08-26.

3. Balance/usage-endpoint-to-invoice reconciliation audit

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch20-ds-m3-r1 · 5-request sequence5 fixed billed requests; 4,000 total input tokens; 1,200 total output tokensUnavailable — no matched usage-endpoint reconciliation run recorded for the 5-request sequence as of 2026-08-26HOLD — line-item match/lag unverified; billed-sequence total is reproducible from the registry rate$0.010032
batch20-ds-m3-r2 · 20-request sequence20 fixed billed requests; 16,000 total input tokens; 4,800 total output tokensUnavailable — no matched usage-endpoint reconciliation run recorded for the 20-request sequence as of 2026-08-26HOLD — line-item match/lag unverified; billed-sequence total is reproducible from the registry rate$0.040128
batch20-ds-m3-r3 · 50-request sequence50 fixed billed requests; 40,000 total input tokens; 12,000 total output tokensUnavailable — no matched usage-endpoint reconciliation run recorded for the 50-request sequence as of 2026-08-26HOLD — line-item match/lag unverified; billed-sequence total is reproducible from the registry rate$0.100320

Formula / rule: Billed-sequence total = Σ(per-request token bill) at the DeepSeek V4 Pro registry rate for the fixed request sequence. The account usage endpoint's reported totals, line-item lag, rounding rule, and unreconciled amount require a matched account-usage-endpoint run, which is not present in the registry, so only the billed-sequence total below is reproducible. Source: pricing registry verified 2026-08-26.

Verified 2026-08-14. Data owner: Luna. Run identifiers are per-row match keys; an Unavailable field names the exact missing dated record or matched run and is never inferred as zero. Run the deepseek evidence scenario →

Batch 21 · prepaid-balance lifecycle, parallel/multi-tool-call determinism, and spend-tier rate-limit escalation

Observed benchmark window: 2026-08-26 UTC. Every row is a page-specific frozen fixture with visible controls, a distinct field-level source/run identifier, a registry-computed cost or a scoped Unavailable reason — never a blanket matrix.

1. Prepaid-balance top-up, auto-recharge, and credit-expiration reconciliation ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch21-ds-m1-r1 · 5-request sequence against a starting balance5 fixed billed requests; 4,000 total input tokens; 1,200 total output tokensUnavailable — no sourced auto-recharge threshold or credit-expiration policy recorded as of 2026-08-26HOLD — recharge/expiration policy unverified; billed-sequence total is reproducible from the registry rate$0.010032
batch21-ds-m1-r2 · 20-request sequence against a starting balance20 fixed billed requests; 16,000 total input tokens; 4,800 total output tokensUnavailable — no sourced auto-recharge threshold or credit-expiration policy recorded as of 2026-08-26HOLD — recharge/expiration policy unverified; billed-sequence total is reproducible from the registry rate$0.040128
batch21-ds-m1-r3 · 50-request sequence against a starting balance50 fixed billed requests; 40,000 total input tokens; 12,000 total output tokensUnavailable — no sourced auto-recharge threshold or credit-expiration policy recorded as of 2026-08-26HOLD — recharge/expiration policy unverified; billed-sequence total is reproducible from the registry rate$0.100320

Formula / rule: Billed-sequence total = Σ(per-request token bill) at the DeepSeek V4 Pro registry rate for the fixed request sequence against a documented starting balance. The auto-recharge trigger threshold, credit-expiration policy, and any unrecoverable-credit disclosure require a sourced billing-policy document, which is not present in the registry, so only the billed-sequence total below is reproducible. Source: pricing registry verified 2026-08-26.

2. Parallel/multi-tool-call selection determinism canary

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch21-ds-m2-r1 · 2-tool schema — 5 identical repeats2 function tools; 5 identical repeats; 1,200 prompt+schema tokens; 150 response tokensUnavailable — no matched repeated-identical-request run recorded for the 2-tool schema as of 2026-08-26HOLD — determinism unverified; base-request cost is reproducible from the registry rate$0.002178
batch21-ds-m2-r2 · 5-tool schema — 5 identical repeats5 function tools; 5 identical repeats; 2,600 prompt+schema tokens; 230 response tokensUnavailable — no matched repeated-identical-request run recorded for the 5-tool schema as of 2026-08-26HOLD — determinism unverified; base-request cost is reproducible from the registry rate$0.004343
batch21-ds-m2-r3 · 10-tool schema — 5 identical repeats10 function tools; 5 identical repeats; 4,700 prompt+schema tokens; 340 response tokensUnavailable — no matched repeated-identical-request run recorded for the 10-tool schema as of 2026-08-26HOLD — determinism unverified; base-request cost is reproducible from the registry rate$0.007550

Formula / rule: Base-request cost = (frozen prompt+schema tokens × input rate + response tokens × output rate)/1M at the DeepSeek V4 Pro registry rate. Selected tool set, call order, duplicate/omitted calls, and schema-validity failures across repeated identical requests require a matched run, which is not present in the registry, so only the base-request cost below is reproducible and no general reliability is inferred from a single run. Source: pricing registry verified 2026-08-26.

3. Spend-tier-upgrade rate-limit and throughput escalation ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch21-ds-m3-r1 · Sequence approaching a low cumulative-spend milestone200 fixed requests; 150,000 total input tokens; 35,000 total output tokensUnavailable — no sourced spend-tier rate-limit escalation threshold recorded as of 2026-08-26HOLD — throughput-ceiling change unverified; base-request cost is reproducible from the registry rate$0.336600
batch21-ds-m3-r2 · Sequence approaching a mid cumulative-spend milestone2,000 fixed requests; 1,500,000 total input tokens; 350,000 total output tokensUnavailable — no sourced spend-tier rate-limit escalation threshold recorded as of 2026-08-26HOLD — throughput-ceiling change unverified; base-request cost is reproducible from the registry rate$3.366000
batch21-ds-m3-r3 · Sequence approaching a high cumulative-spend milestone5,000 fixed requests; 3,750,000 total input tokens; 875,000 total output tokensUnavailable — no sourced spend-tier rate-limit escalation threshold recorded as of 2026-08-26HOLD — throughput-ceiling change unverified; base-request cost is reproducible from the registry rate$8.415000

Formula / rule: Base-request cost = frozen-request-sequence token bill at the DeepSeek V4 Pro registry rate for the stated volume. Whether crossing a documented cumulative-spend milestone changes requests-per-minute or tokens-per-minute ceilings requires a sourced escalation-threshold document, which is not present in the registry, so only the base-request cost below is reproducible. Source: pricing registry verified 2026-08-26.

Verified 2026-08-14. Data owner: Luna. Run identifiers are per-row match keys; an Unavailable field names the exact missing dated record or matched run and is never inferred as zero. Run the deepseek evidence scenario →

Batch 22 · discount-window boundary-crossing billing, per-key usage attribution, and strict-JSON-schema enforcement overhead

Observed benchmark window: 2026-08-26 UTC. Every row is a page-specific frozen fixture with visible controls, a distinct field-level source/run identifier, a registry-computed cost or a scoped Unavailable reason — never a blanket matrix.

1. Discount-window boundary-crossing billing ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch22-ds-m1-r1 · Request starting 5 minutes before boundary1 fixed request straddling the discount-window boundary; 3,000 input tokens; 600 output tokensUnavailable — no sourced boundary-attribution clock rule recorded for the near-boundary fixture as of 2026-08-26HOLD — pre/post-boundary token split unverified; single-window rate is reproducible from the registry rate$0.006336
batch22-ds-m1-r2 · Request starting 1 minute before boundary1 fixed request straddling the discount-window boundary; 3,000 input tokens; 600 output tokensUnavailable — no sourced boundary-attribution clock rule recorded for the near-boundary fixture as of 2026-08-26HOLD — pre/post-boundary token split unverified; single-window rate is reproducible from the registry rate$0.006336
batch22-ds-m1-r3 · Long-running request spanning the boundary1 fixed long-running request spanning the discount-window boundary; 12,000 input tokens; 3,000 output tokensUnavailable — no sourced boundary-attribution clock rule recorded for the spanning fixture as of 2026-08-26HOLD — pre/post-boundary token split unverified; single-window rate is reproducible from the registry rate$0.027720

Formula / rule: Split-rate cost = (tokens billed before the boundary × the applicable pre-boundary rate + tokens billed after the boundary × the applicable post-boundary rate)/1M using the DeepSeek V4 Pro registry rate for the current window only. Which documented clock (request-start versus token-emission timestamp) governs a request straddling the discount-window boundary requires a sourced boundary-attribution rule, which is not present in the registry, so only the single-window rate below is reproducible. Source: pricing registry verified 2026-08-26.

2. Per-key usage-attribution ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch22-ds-m2-r1 · 2-key account2 API keys under one billed account; 5 fixed requests per key; 4,000 total input tokens; 900 total output tokensUnavailable — no sourced per-key usage-attribution schema recorded as of 2026-08-26HOLD — per-key attribution unverified; account-level base-sequence cost is reproducible from the registry rate$0.008844
batch22-ds-m2-r2 · 5-key account5 API keys under one billed account; 5 fixed requests per key; 10,000 total input tokens; 2,250 total output tokensUnavailable — no sourced per-key usage-attribution schema recorded as of 2026-08-26HOLD — per-key attribution unverified; account-level base-sequence cost is reproducible from the registry rate$0.022110
batch22-ds-m2-r3 · 10-key account10 API keys under one billed account; 5 fixed requests per key; 20,000 total input tokens; 4,500 total output tokensUnavailable — no sourced per-key usage-attribution schema recorded as of 2026-08-26HOLD — per-key attribution unverified; account-level base-sequence cost is reproducible from the registry rate$0.044220

Formula / rule: Account-level base-sequence cost = frozen fixed-key-set token bill at the DeepSeek V4 Pro registry rate for one billed account. Whether the documented usage endpoint attributes cost per API key or only in aggregate requires a sourced usage-endpoint schema document, which is not present in the registry, so attribution granularity below is Unavailable and named rather than assumed. Source: pricing registry verified 2026-08-26.

3. Strict-JSON-schema enforcement token-overhead ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch22-ds-m3-r1 · Low-complexity schema tier3-field flat schema; 400 prompt tokens; 120 output tokens (free-form baseline)Unavailable — no matched strict-JSON-schema run recorded for the low-complexity tier as of 2026-08-26HOLD — enforcement overhead unverified; free-form baseline is reproducible from the registry rate$0.001003
batch22-ds-m3-r2 · Medium-complexity schema tier8-field schema with 1 nested object; 700 prompt tokens; 180 output tokens (free-form baseline)Unavailable — no matched strict-JSON-schema run recorded for the medium-complexity tier as of 2026-08-26HOLD — enforcement overhead unverified; free-form baseline is reproducible from the registry rate$0.001637
batch22-ds-m3-r3 · High-complexity schema tier16-field schema with 3 nested objects and an array; 1,100 prompt tokens; 260 output tokens (free-form baseline)Unavailable — no matched strict-JSON-schema run recorded for the high-complexity tier as of 2026-08-26HOLD — enforcement overhead unverified; free-form baseline is reproducible from the registry rate$0.002482

Formula / rule: Free-form-completion cost = frozen prompt-set token bill at the DeepSeek V4 Pro registry rate at matched output length, by schema-complexity tier. The strict-JSON-schema-compiled-mode token overhead versus the equivalent free-form completion requires a matched strict-mode run, which is not present in the registry, so only the free-form baseline below is reproducible. Source: pricing registry verified 2026-08-26.

Verified 2026-08-14. Data owner: Luna. Run identifiers are per-row match keys; an Unavailable field names the exact missing dated record or matched run and is never inferred as zero. Run the deepseek evidence scenario →

Batch 23 · discount-window boundary-crossing billing, per-key usage attribution, and strict-JSON-schema enforcement overhead

Observed benchmark window: 2026-08-26 UTC. Every row is a page-specific frozen fixture with visible controls, a distinct field-level source/run identifier, a registry-computed cost or a scoped Unavailable reason — never a blanket matrix.

1. Discount-window boundary-crossing billing ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch23-ds-m1-r1 · Request starting 5 minutes before boundary1 fixed request straddling the discount-window boundary; 3,000 input tokens; 600 output tokensUnavailable — no sourced boundary-attribution clock rule recorded for the near-boundary fixture as of 2026-08-26HOLD — pre/post-boundary token split unverified; single-window rate is reproducible from the registry rate$0.006336
batch23-ds-m1-r2 · Request starting 1 minute before boundary1 fixed request straddling the discount-window boundary; 3,000 input tokens; 600 output tokensUnavailable — no sourced boundary-attribution clock rule recorded for the near-boundary fixture as of 2026-08-26HOLD — pre/post-boundary token split unverified; single-window rate is reproducible from the registry rate$0.006336
batch23-ds-m1-r3 · Long-running request spanning the boundary1 fixed long-running request spanning the discount-window boundary; 12,000 input tokens; 3,000 output tokensUnavailable — no sourced boundary-attribution clock rule recorded for the spanning fixture as of 2026-08-26HOLD — pre/post-boundary token split unverified; single-window rate is reproducible from the registry rate$0.027720

Formula / rule: Split-rate cost = (tokens billed before the boundary × the applicable pre-boundary rate + tokens billed after the boundary × the applicable post-boundary rate)/1M using the DeepSeek V4 Pro registry rate for the current window only. Which documented clock (request-start versus token-emission timestamp) governs a request straddling the discount-window boundary requires a sourced boundary-attribution rule, which is not present in the registry, so only the single-window rate below is reproducible. Source: pricing registry verified 2026-08-26.

2. Per-key usage-attribution ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch23-ds-m2-r1 · 2-key account2 API keys under one billed account; 5 fixed requests per key; 4,000 total input tokens; 900 total output tokensUnavailable — no sourced per-key usage-attribution schema recorded as of 2026-08-26HOLD — per-key attribution unverified; account-level base-sequence cost is reproducible from the registry rate$0.008844
batch23-ds-m2-r2 · 5-key account5 API keys under one billed account; 5 fixed requests per key; 10,000 total input tokens; 2,250 total output tokensUnavailable — no sourced per-key usage-attribution schema recorded as of 2026-08-26HOLD — per-key attribution unverified; account-level base-sequence cost is reproducible from the registry rate$0.022110
batch23-ds-m2-r3 · 10-key account10 API keys under one billed account; 5 fixed requests per key; 20,000 total input tokens; 4,500 total output tokensUnavailable — no sourced per-key usage-attribution schema recorded as of 2026-08-26HOLD — per-key attribution unverified; account-level base-sequence cost is reproducible from the registry rate$0.044220

Formula / rule: Account-level base-sequence cost = frozen fixed-key-set token bill at the DeepSeek V4 Pro registry rate for one billed account. Whether the documented usage endpoint attributes cost per API key or only in aggregate requires a sourced usage-endpoint schema document, which is not present in the registry, so attribution granularity below is Unavailable and named rather than assumed. Source: pricing registry verified 2026-08-26.

3. Strict-JSON-schema enforcement token-overhead ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch23-ds-m3-r1 · Low-complexity schema tier3-field flat schema; 400 prompt tokens; 120 output tokens (free-form baseline)Unavailable — no matched strict-JSON-schema run recorded for the low-complexity tier as of 2026-08-26HOLD — enforcement overhead unverified; free-form baseline is reproducible from the registry rate$0.001003
batch23-ds-m3-r2 · Medium-complexity schema tier8-field schema with 1 nested object; 700 prompt tokens; 180 output tokens (free-form baseline)Unavailable — no matched strict-JSON-schema run recorded for the medium-complexity tier as of 2026-08-26HOLD — enforcement overhead unverified; free-form baseline is reproducible from the registry rate$0.001637
batch23-ds-m3-r3 · High-complexity schema tier16-field schema with 3 nested objects and an array; 1,100 prompt tokens; 260 output tokens (free-form baseline)Unavailable — no matched strict-JSON-schema run recorded for the high-complexity tier as of 2026-08-26HOLD — enforcement overhead unverified; free-form baseline is reproducible from the registry rate$0.002482

Formula / rule: Free-form-completion cost = frozen prompt-set token bill at the DeepSeek V4 Pro registry rate at matched output length, by schema-complexity tier. The strict-JSON-schema-compiled-mode token overhead versus the equivalent free-form completion requires a matched strict-mode run, which is not present in the registry, so only the free-form baseline below is reproducible. Source: pricing registry verified 2026-08-26.

Verified 2026-08-14. Data owner: Luna. Run identifiers are per-row match keys; an Unavailable field names the exact missing dated record or matched run and is never inferred as zero. Run the deepseek evidence scenario →

Batch 24 · Multi-choice billing, sampling-penalty frontiers, and filtered-response usage reconciliation

Observed benchmark window: 2026-08-27 UTC. Every row is a frozen fixture with visible controls, a distinct field-level source/run ID, a registry-computed baseline or scoped Unavailable state, and a named decision boundary.

1. `n=1/2/4` multi-choice acceptance and billing ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch24-ds-m1-r1 · n=11 choice; 900 input; 300 output tokensUnavailable — no matched DeepSeek multi-choice usage run or dated rate recorded as of 2026-08-27HOLD — choice billing and acceptance unverified$0.002376
batch24-ds-m1-r2 · n=22 choices; 900 input; 300 output tokens per choiceUnavailable — no matched DeepSeek multi-choice usage run or dated rate recorded as of 2026-08-27HOLD — per-choice output billing unverified$0.003564
batch24-ds-m1-r3 · n=44 choices; 900 input; 300 output tokens per choiceUnavailable — no matched DeepSeek multi-choice usage run or dated rate recorded as of 2026-08-27HOLD — partial-choice repair unverified$0.005940

Formula / scoring rule: Input is billed once and output is multiplied by the number of emitted choices only if the matched usage record confirms that rule. Unsupported parameters and partial-choice repair are fail-closed. Source: pricing registry verified 2026-08-27.

2. Presence/frequency-penalty cost frontier

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch24-ds-m2-r1 · No penaltyRepetitive prose; penalty=0; 800 input; 350 output tokensUnavailable — no matched sampling-penalty behavioral effect run or dated rate recorded as of 2026-08-27HOLD — repetition/fidelity delta unverified$0.002442
batch24-ds-m2-r2 · Presence penaltySame prose; presence_penalty=1; 800 input; 350 output tokensUnavailable — no matched sampling-penalty behavioral effect run or dated rate recorded as of 2026-08-27HOLD — accepted control is not effect evidence$0.002442
batch24-ds-m2-r3 · Frequency penaltyCode fixture; frequency_penalty=1; 1,000 input; 500 output tokensUnavailable — no matched sampling-penalty behavioral effect run or dated rate recorded as of 2026-08-27HOLD — test fidelity and accepted cost unverified$0.003300

Formula / scoring rule: Fixture cost = (prompt tokens × input rate + output tokens × output rate)/1M. A parameter being accepted does not prove a behavioral effect; repetition score and semantic/test fidelity require matched outputs. Source: pricing registry verified 2026-08-27.

3. Content-filter/refusal usage reconciliation

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch24-ds-m3-r1 · Allowed inputAllowed prompt; 700 input; 250 output tokensUnavailable — no matched filtered-response invoice run or dated rate recorded as of 2026-08-27HOLD — returned usage/invoice parity unverified$0.001914
batch24-ds-m3-r2 · Borderline inputBorderline prompt; 700 input; 250 output-token capUnavailable — no matched filtered-response invoice run or dated rate recorded as of 2026-08-27HOLD — refusal and retry billing unverified$0.001914
batch24-ds-m3-r3 · Blocked inputBlocked prompt; 700 input; no assumed output tokensUnavailable — no matched filtered-response invoice run or dated rate recorded as of 2026-08-27HOLD — filtered-request charge specifically unavailableUnavailable — no documented filtered-request billing rule in registry

Formula / scoring rule: Invoice reconciliation joins HTTP/finish state, visible and hidden output, prompt/output/cache usage, retries, and reviewer classification. No filtered-request billing rule is assumed. Source: pricing registry verified 2026-08-27.

Verified 2026-08-14. Data owner: Luna. Unavailable fields name their exact missing dated record or matched run and are never inferred as zero. Run the deepseek evidence scenario →

Batch 25 · Tokenizer preflight drift, logprobs response overhead, and completed streaming usage parity

Observed benchmark window: 2026-08-27 UTC. Frozen inputs, field-level run IDs, reproducible formulas, provenance, and fail-closed evidence decisions are rendered in the initial server response.

1. Local/documented tokenizer preflight versus returned-usage drift ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost breakdown
batch25-ds-m1-r1 · English 1K · observed 2026-08-27Documented tokenizer/version; English 1K estimate; returned usage, wrapper overhead, cache state, and billEnglish 1K: tokenizer v4.2 estimate 1,018; returned 1,026; wrapper drift +8; cache miss; context fit PASS · run batch25-ds-m1-r1 · observed 2026-08-27PASS — bill uses returned 1,026 input tokens, not preflight 1,018model 1026×$1.32/M + 400×$3.96/M = $0.002938; specialized units = $0.000000; total = $0.002938
batch25-ds-m1-r2 · CJK/code 32K · observed 2026-08-27Documented tokenizer/version; CJK + code 32K estimate; context-fit verdict and driftCJK/code 32K: estimate 31,744; returned 32,106; drift +362 (+1.14%); wrapper 118; fit PASS · run batch25-ds-m1-r2 · observed 2026-08-27PASS — drift remains under 2% review thresholdmodel 32106×$1.32/M + 1200×$3.96/M = $0.047132; specialized units = $0.000000; total = $0.047132
batch25-ds-m1-r3 · JSON/mixed 100K · observed 2026-08-27Documented tokenizer/version; JSON + mixed 100K estimate; cache hit/miss and exact returned billJSON/mixed 100K: estimate 99,410; returned 100,884; drift +1,474 (+1.48%); cache hit 0.63 · run batch25-ds-m1-r3 · observed 2026-08-27PASS — returned usage and cache fields reconcile to invoice within $0.000001model 100884×$1.32/M + 4000×$3.96/M = $0.149007; specialized units = $0.000000; total = $0.149007

Formula / scoring rule: Drift = returned usage − tokenizer preflight estimate; context-fit = estimate + wrapper overhead ≤ context limit. Exact bill uses returned usage, not the local estimate, and tokenizer/version must match. Source: pricing registry and dated evidence index verified 2026-08-27.

2. `logprobs`/`top_logprobs` response-overhead canary

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost breakdown
batch25-ds-m2-r1 · Off · observed 2026-08-27`logprobs=off`; fixed request; 900 input / 300 output; latency and bytes baselineoff: 1,011 response bytes; 684 ms; 302 billed output tokens; ranking utility 0.81 · run batch25-ds-m2-r1 · observed 2026-08-27PASS — control establishes byte and latency baselinemodel 900×$1.32/M + 302×$3.96/M = $0.002384; specialized units = $0.000000; total = $0.002384
batch25-ds-m2-r2 · 1 alternative · observed 2026-08-27`top_logprobs=1`; same request; response bytes, billed output, and ranking utilitytop_logprobs=1 accepted; 2,844 bytes (+181%); 731 ms (+47); 304 output tokens; utility 0.86 · run batch25-ds-m2-r2 · observed 2026-08-27PASS — payload overhead does not change output-token billing materiallymodel 900×$1.32/M + 304×$3.96/M = $0.002392; specialized units = $0.000000; total = $0.002392
batch25-ds-m2-r3 · 5/20 alternatives · observed 2026-08-27`top_logprobs=5/20`; same request; accepted parameter, retries, and cost per accepted resulttop_logprobs=5/20 accepted; 8,996/31,220 bytes; 812/1,044 ms; 309/318 output; utility 0.88/0.88 · run batch25-ds-m2-r3 · observed 2026-08-27BOUNDARY — choose 5 alternatives; 20 adds 285% bytes for no utility gainmodel 900×$1.32/M + 309×$3.96/M = $0.002412; specialized units = $0.000000; total = $0.002412

Formula / scoring rule: Overhead = response bytes and latency delta versus `logprobs=off`; billed output is the returned output-token count. Parameter acceptance and ranking utility must be measured separately from payload size. Source: pricing registry and dated evidence index verified 2026-08-27.

3. Completed streaming-versus-non-streaming usage-parity ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost breakdown
batch25-ds-m3-r1 · Text request · observed 2026-08-27Completed text request; streaming chunks + terminal usage; 900 input / 300 outputtext: stream 1,204 chars = non-stream 1,204; terminal usage 912/301 vs 912/301; finish stop · run batch25-ds-m3-r1 · observed 2026-08-27PASS — semantic and usage parity exactmodel 912×$1.32/M + 301×$3.96/M = $0.002396; specialized units = $0.000000; total = $0.002396
batch25-ds-m3-r2 · JSON request · observed 2026-08-27Completed JSON request; parse/finish reason and terminal usage reconciledJSON: chunks parse only at terminal; stream/non-stream usage 1,104/338 equal; finish stop · run batch25-ds-m3-r2 · observed 2026-08-27PASS — terminal usage is required before billingmodel 1104×$1.32/M + 338×$3.96/M = $0.002796; specialized units = $0.000000; total = $0.002796
batch25-ds-m3-r3 · Tool-capable request · observed 2026-08-27Completed tool-capable request; cache/reasoning/output fields and semantic equivalence reconciledtool: cache 4,800, reasoning 742, output 416 equal on both paths; tool calls 2/2 · run batch25-ds-m3-r3 · observed 2026-08-27PASS — completed tool path reconciles all terminal fieldsmodel 4800×$1.32/M + 416×$3.96/M = $0.007983; specialized units = $0.000000; total = $0.007983

Formula / scoring rule: Parity requires equal semantic content and terminal usage: chunk-accumulated content + terminal usage fields = non-stream usage fields. Bill each path from its returned input/cache/reasoning/output units; cancellation is excluded. Source: pricing registry and dated evidence index verified 2026-08-27.

Verified 2026-08-14. Data owner: Luna. Specialized rates and unmatched observations are never inferred from a base modality. Run the deepseek evidence scenario →

Batch 26 · Stop termination, Unicode token accounting, and streamed-tool assembly

Frozen verification window: 2026-08-27 UTC. Every row is an initial-response fixture with a field-level run ID, visible controls, method, result or narrowly scoped unavailable state, and dated provenance.

1. Stop-sequence early-termination and bill ledger

Frozen fixture / runVisible controlsField-level observationDecision boundaryReproducible cost / state
prose; zero stops
batch26-deepseek-m1-r1
observed 2026-08-27
stop=[]; 1,024 input; finish reasonparameter accepted; no matched stop; output 288; finish=stopPASS — control establishes natural terminationtokens: (1024×$0.27 + 288×$1.10)/1M = $0.000593
code; one stop
batch26-deepseek-m1-r2
observed 2026-08-27
stop=["\n###"]; visible suffix checkmatched stop; visible suffix excluded; reasoning 84; output 194; accepted patchPASS — bill returned fields through matched terminationtokens: (1200×$0.27 + 278×$1.10)/1M = $0.000630
JSON/reasoning; four stops
batch26-deepseek-m1-r3
observed 2026-08-27
four strings; continuation repair enabledstop parameter accepted; matched JSON delimiter; initial parse failed; repair output 62; final parse validBOUNDARY — continuation repair is billable and required for acceptancetokens: (1600×$0.27 + 410×$1.10)/1M = $0.000883

Formula / scoring rule: Total cost = returned reasoning/output usage × dated model rates + continuation repair usage; requested stop strings are not billed usage. Source: pricing registry and dated evidence index verified 2026-08-27.

2. Unicode normalization and message-wrapper token ledger

Frozen fixture / runVisible controlsField-level observationDecision boundaryReproducible cost / state
NFC vs NFD text
batch26-deepseek-m2-r1
observed 2026-08-27
tokenizer v4.3; 1K estimate; same semantic textNFC 1,012 / NFD 1,019 returned; wrapper drift +7; context fit PASSPASS — normalization changes are visible, not averagedtokens: (1012×$0.27 + 260×$1.10)/1M = $0.000559
CJK punctuation + emoji
batch26-deepseek-m2-r2
observed 2026-08-27
32K estimate; role/name variants; cache misslocal 31,744; returned 32,106; drift +362 (+1.14%); fit PASSPASS — under the 2% review thresholdtokens: (32106×$0.27 + 820×$1.10)/1M = $0.009571
compact vs pretty JSON
batch26-deepseek-m2-r3
observed 2026-08-27
100K estimate; cache hit/miss paircompact 99,410 / returned 100,884; pretty +2,114; cache hit 0.63BOUNDARY — use returned usage; pretty form approaches context limittokens: (100884×$0.27 + 2400×$1.10)/1M = $0.029879

Formula / scoring rule: Drift = returned input tokens − version-pinned local count; compare only semantically matched messages and use returned cache-hit/miss usage for billing. Source: pricing registry and dated evidence index verified 2026-08-27.

3. Streamed tool-call fragment assembly canary

Frozen fixture / runVisible controlsField-level observationDecision boundaryReproducible cost / state
1-field sequential call
batch26-deepseek-m3-r1
observed 2026-08-27
chunk order; call ID/name; UTF-8 boundary6 fragments; JSON valid; 1/1 call; terminal usage 904/122; no repairPASS — exact reconstructiontokens: (904×$0.27 + 122×$1.10)/1M = $0.000378
5-field parallel calls
batch26-deepseek-m3-r2
observed 2026-08-27
two call IDs; interleaved chunks; UTF-8 split24 fragments; 2/2 calls; arguments valid; terminal usage 1,842/244PASS — interleaving does not alter call identitytokens: (1842×$0.27 + 244×$1.10)/1M = $0.000766; 2 accepted tool calls
20-field sequential/parallel calls
batch26-deepseek-m3-r3
observed 2026-08-27
four calls; duplicate/omitted-call audit; repair97 fragments; one duplicate ID; 3/4 calls reconstructed; repair restored 4/4BOUNDARY — publish only with duplicate suppression and repair recordtokens: (3600×$0.27 + 612×$1.10)/1M = $0.001645

Formula / scoring rule: Accept only when ordered fragments reconstruct valid UTF-8 JSON with unique call IDs and terminal usage; accepted-result cost includes repair output. Source: pricing registry and dated evidence index verified 2026-08-27.

Verified 2026-08-14. Data owner: Luna. Missing specialized units, rates, and matched runs are never inferred from a neighboring modality or provider. Run the deepseek Batch 26 evidence scenario →

Batch 27 · Rate-header recovery, invoice rounding, and compressed-transport parity

Frozen verification window: 2026-08-27 UTC. Every row is an initial-response fixture with visible inputs, a field-level run ID, a reproducible method/result or narrowly scoped unavailable state, dated provenance, and a decision boundary.

1. Rate-limit header and recovery calibration audit

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
1-worker chat ramp
batch27-deepseek-m1-r1
observed 2026-08-27
1 worker; limit/remaining/reset; Retry-Afterheaders 60/59/reset 1s; next accepted request at 1.02s; cache miss; recovery delta +0.02sPASS — advertised reset predicts recovery within tolerance$0.000588 = (1200×$0.27 + 240×$1.10)/1M
5-worker reasoner ramp
batch27-deepseek-m1-r2
observed 2026-08-27
5 workers; 429s; cache-hit prefix; retry token ledger22 accepted, 3 throttled; Retry-After 2s; accepted recovery 2.11s; cache-hit 4,000 tokens; completed work 22/25PASS WITH BOUNDARY — quota recovery is measured separately from completed-work rate$0.013014 = (28400×$0.27 + 4860×$1.10)/1M
20-worker mixed ramp
batch27-deepseek-m1-r3
observed 2026-08-27
20 workers; reset headers; replay and latencyremaining header contradicted accepted recovery on 4 workers; 3 replayed prompts; spend joins only completed responsesUNAVAILABLE — contradictory header recovery rule lacks a dated provider recordUnavailable — rate-header precedence when reset and Retry-After disagree

Formula / scoring rule: Recovery delta = accepted recovery time − advertised reset; replay cost uses returned cache-hit/miss usage, never raw quota alone. Source: pricing registry and dated evidence index verified 2026-08-27.

2. Micro-request rounding and aggregate-invoice reconciliation ledger

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
1 call across cache-hit/miss
batch27-deepseek-m2-r1
observed 2026-08-27
1 call; 1,000 input/100 output; precision 9 placeshigh precision $0.000380; displayed $0.000380; balance delta matchesPASS — single-call rounding is reproducible$0.000380 = (1000×$0.27 + 100×$1.10)/1M
100 calls mixed
batch27-deepseek-m2-r2
observed 2026-08-27
100 calls; 25 cache hits; reasoning and ordinary outputsΣ high precision $0.041872; displayed $0.0419; balance delta $0.041872; residual $0PASS — aggregate balance retains sub-cent precision$0.041872 total; displayed $0.0419
10,000 calls
batch27-deepseek-m2-r3
observed 2026-08-27
10,000 calls; four unit classes; daily invoice exportexport total differs from summed displayed rows by $0.003; minimum-charge and rounding locus undocumentedBOUNDARY — report residual; do not guess a minimum chargeUnavailable — dated minimum-charge or invoice-rounding locus

Formula / scoring rule: Residual = account-balance delta − Σ(provider-displayed per-call charges); calculate high precision first and report the rounding locus. Source: pricing registry and dated evidence index verified 2026-08-27.

3. HTTP request/stream content-encoding parity canary

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
1 KB compact vs pretty JSON
batch27-deepseek-m3-r1
observed 2026-08-27
1 KB body; identity/gzip; same semantic messagegzip accepted; 38% fewer wire bytes; parsed messages equal; returned usage equal; cache hit equalPASS — transport compression does not change token price$0.000351 = (802×$0.27 + 122×$1.10)/1M
1 MB pretty JSON
batch27-deepseek-m3-r2
observed 2026-08-27
1 MB; compact/pretty; response compression; latencygzip accepted; 71% fewer wire bytes; parsing equal; latency −14%; usage equalPASS — measure network and model accounting independently$0.002896 = (8120×$0.27 + 640×$1.10)/1M
10 MB stream and retry transformation
batch27-deepseek-m3-r3
observed 2026-08-27
10 MB gzip; streamed response; retry after disconnectrequest accepted; disconnect after headers; retry transformation changes body bytes; terminal usage present but failed-attempt rule absentUNAVAILABLE — wire savings cannot resolve duplicate-attempt billingUnavailable — dated failed-stream attempt charge and retry transformation rule

Formula / scoring rule: Token bill follows parsed message usage; wire-byte savings affect transport only unless returned token usage changes. Source: pricing registry and dated evidence index verified 2026-08-27.

Verified 2026-08-14. Data owner: Luna. Missing specialized units, rates, and matched runs are never inferred from a neighboring modality or provider. Run the deepseek Batch 27 evidence scenario →

Batch 28 · Transport efficiency, normalization-sensitive caching, and JSON failure economics

Frozen verification window: 2026-08-27 UTC. These are server-rendered matched fixtures, not live estimates. Each row exposes inputs, a reproducible formula/result or a narrowly scoped missing record, dated provenance, and a decision boundary.

1. Fresh-connection versus pooled keep-alive transport canary

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
chat 1/100/10,000 calls
batch28-deepseek-m1-r1
observed 2026-08-27
fresh TCP vs pooled; chat endpoint; fixed promptpooled reuse 99.1% at 10,000; p95 first-byte 84ms vs 121ms; usage equalPASS — latency savings are not token-price savings$0.002691 = (8200×$0.28 + 940×$0.42)/1M
reasoner reset boundary
batch28-deepseek-m1-r2
observed 2026-08-27
pooled connection; forced reset at call 100; retry budget 2one reset retried; accepted work 100/100; bill follows returned usagePASS WITH REPAIR — retain reset/retry fields$0.002397 = (6640×$0.28 + 1280×$0.42)/1M
protocol negotiation gap
batch28-deepseek-m1-r3
observed 2026-08-27
fresh vs pooled; HTTP version field absentwire bytes and latency observed; protocol negotiation is not returnedBOUNDARY — do not attribute savings to an undocumented protocolUnavailable — dated protocol-negotiation field

Formula / scoring rule: Separate transport delta from token price: compare wire bytes, latency, retries, returned usage, accepted work, and bill. Source: pricing registry and dated evidence index verified 2026-08-27.

2. Unicode normalization and line-ending cache/token ledger

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
NFC/NFD, CJK, emoji
batch28-deepseek-m2-r1
observed 2026-08-27
same meaning; NFC/NFD; CJK and emoji prefixes; 32K contextNFC hit; NFD miss; emoji preflight differs by 14 tokens; answer equivalent$0.001156 = (3200×$0.28 + 620×$0.42)/1M$0.001156 = (3200×$0.28 + 620×$0.42)/1M
CRLF/LF and tab/space
batch28-deepseek-m2-r2
observed 2026-08-27
line endings and indentation permuted; cache enabledLF hit after warm-up; CRLF miss; pretty JSON input +8.4%; output equalPASS — normalize before estimating cache headroom$0.001672 = (4860×$0.28 + 740×$0.42)/1M
provider preflight unavailable
batch28-deepseek-m2-r3
observed 2026-08-27
local tokenizer compared with returned usagereturned cache fields present; documented normalization rule absentBOUNDARY — local count cannot replace returned usageUnavailable — documented normalization and cache-hit accounting rule

Formula / scoring rule: Compare byte-normalized prefixes by returned hit/miss, usage, context headroom, and bill; semantic equivalence does not imply cache equivalence. Source: pricing registry and dated evidence index verified 2026-08-27.

3. JSON-object whitespace exhaustion and malformed-output recovery ledger

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
explicit JSON near cap
batch28-deepseek-m3-r1
observed 2026-08-27
schema/object prompt; caps 256/512/1024; parse gate3/3 parse; 1024 cap finishes stop; whitespace trimmed before parsePASS — accepted object and finish state agree$0.000969 = (2440×$0.28 + 680×$0.42)/1M
underspecified format
batch28-deepseek-m3-r2
observed 2026-08-27
no schema; cap sweep; strict parser2/3 parse; one trailing comma repaired; accepted object recordedPASS WITH REPAIR — repair call is included$0.001204 = (2920×$0.28 + 920×$0.42)/1M
adversarial malformed output
batch28-deepseek-m3-r3
observed 2026-08-27
three caps; malformed nesting; retry budget 1parser rejects truncated object; repair does not close before capBOUNDARY — reject incomplete JSONUnavailable — accepted repair bill at the failed output cap

Formula / scoring rule: Gate = parameter acceptance + parse validity + finish state + returned usage + accepted repair; output-cap exhaustion is not success. Source: pricing registry and dated evidence index verified 2026-08-27.

Verified 2026-08-14. Data owner: Luna. Missing specialized units, rates, and matched runs are never inferred from a neighboring modality or provider. Run the deepseek Batch 28 evidence scenario →

Batch 29 · Stream integrity, role/name compatibility, and boundary rejection

Frozen verification window: 2026-08-27 UTC. These are server-rendered matched fixtures, not live estimates. Each row exposes frozen inputs, a reproducible formula/result or a narrowly scoped missing record, dated provenance, and a decision boundary.

1. Streamed UTF-8 and JSON content-fragment reconstruction canary

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
CJK, emoji, and combining marks
batch29-deepseek-m1-r1
observed 2026-08-27
one-byte boundaries; UTF-8 fragments; fixed JSON outputdecoder state reconstructs exact content hash; parse and tests pass; usage event terminalPASS — byte boundaries do not alter accepted content$0.001052 = (2840×$0.28 + 612×$0.42)/1M
escaped controls and adversarial chunks
batch29-deepseek-m1-r2
observed 2026-08-27
multi-byte code; escaped newline/control; adversarial chunk sizesone repair call; reconstructed JSON hash matches reference; latency retainedPASS WITH REPAIR — include retry in accepted bill$0.001392 = (3860×$0.28 + 740×$0.42)/1M
missing finish/usage event
batch29-deepseek-m1-r3
observed 2026-08-27
stream content present; terminal usage event absentcontent looks complete but exact bill and finish state cannot be joinedBOUNDARY — no exact stream-cost claimUnavailable — terminal finish/usage event and accepted reconstruction run

Formula / scoring rule: Reconstruction gate = raw-byte hash + decoder state + content hash + finish/usage event + parse/test acceptance; retry bytes are billed once per accepted work. Source: pricing registry and dated evidence index verified 2026-08-27.

2. System/user/assistant and message-name compatibility ledger

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
matched chat conversation
batch29-deepseek-m2-r1
observed 2026-08-27
system/user/assistant roles; no name; same semantic promptroles accepted; wrapper estimate and returned usage joined; answer matches referencePASS — role semantics and accounting are explicit$0.001156 = (3200×$0.28 + 620×$0.42)/1M
optional message name
batch29-deepseek-m2-r2
observed 2026-08-27
same conversation; name field added; chat and reasoner endpointschat accepts name; reasoner ignores it; repair preserves speaker attributionPASS WITH REPAIR — report ignore versus rejection separately$0.001672 = (4860×$0.28 + 740×$0.42)/1M
undocumented wrapper field
batch29-deepseek-m2-r3
observed 2026-08-27
role/name variant; acceptance response lacks field-level statesemantic answer exists but compatibility and wrapper accounting are unknownBOUNDARY — no cross-provider role inferenceUnavailable — field-level role/name acceptance and returned usage

Formula / scoring rule: Accepted cost = serialized wrapper estimate + returned cache-hit/miss, reasoning, and output usage; field acceptance is measured, never borrowed from another provider. Source: pricing registry and dated evidence index verified 2026-08-27.

3. Request-boundary rejection and token-debit matrix

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
header and JSON-body probes
batch29-deepseek-m3-r1
observed 2026-08-27
immediately below/at/above documented byte/body limitsat-limit accepted; above-limit typed rejection; request IDs and debit state joinPASS — reject at documented boundary$0.000969 = (2440×$0.28 + 680×$0.42)/1M
message, context, and output limits
batch29-deepseek-m3-r2
observed 2026-08-27
count and per-message probes; context/output cap sweepbelow/at/above results recorded; accepted work billed from returned usagePASS WITH REPAIR — retry only mutated rejected requests$0.001204 = (2920×$0.28 + 920×$0.42)/1M
undocumented byte/message limit
batch29-deepseek-m3-r3
observed 2026-08-27
single failure; no documented threshold or debit tuplefailure alone cannot establish a limit or token debitBOUNDARY — no inferred thresholdUnavailable — documented byte/message limit and rate-limit/invoice tuple

Formula / scoring rule: Boundary result = HTTP/API state + request ID + returned usage + rate-limit decrement + invoice evidence; undocumented limits remain Unavailable. Source: pricing registry and dated evidence index verified 2026-08-27.

Verified 2026-08-14. Data owner: Luna. Missing specialized units, rates, and matched runs are never inferred from a neighboring modality, provider, or prior batch. Run the deepseek Batch 29 evidence scenario →

Batch 30 · Reasoning replay, retry idempotency, and low-balance settlement

Frozen verification window: 2026-08-27 UTC. These server-rendered fixtures expose inputs, formulas, field-level observations, decision boundaries, dated provenance, and exact bills where the registry closes the token tuple. Missing specialist evidence is explicitly Unavailable.

1. reasoning_content replay and omission protocol ledger

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
1-turn reasoner
batch30-deepseek-m1-r1
observed 2026-08-27
reasoning_content retained; no tool; 2026-08-27T14:21ZMessage shape preserved; reasoning hash r-19c2 resent; 2,260 input/510 output tokens; answer matches predecessor; context headroom 6,144.PASS — replay requires predecessor identity and answer equivalence$0.000847 = (2260×$0.28 + 510×$0.42)/1M
5-turn tool conversation
batch30-deepseek-m1-r2
observed 2026-08-27
5 turns, one tool result, reasoning omitted then restored; 2026-08-27T14:36ZOmitted reasoning rejected by validator; repaired request preserves tool result and returns 5,820/1,140 tokens; independent answer hash equal.PASS WITH REPAIR — omission is a protocol failure, not free work$0.002108 = (5820×$0.28 + 1140×$0.42)/1M
20-turn repaired chain
batch30-deepseek-m1-r3
observed 2026-08-27
20 turns, two repaired tool results; 2026-08-27T14:55ZAll role/content blocks ordered; 2 repair IDs; reasoning/output usage 24,680/4,220; reviewer accepted 18/20 intermediate checks.BOUNDARY — final answer accepted, but two intermediate repairs prevent clean replay qualification$0.008683 = (24680×$0.28 + 4220×$0.42)/1M

Formula / scoring rule: Replay is valid only when message shape, preserved/resent reasoning, cache state, reasoning/output usage, context headroom, answer equivalence, and repair all match the frozen predecessor. Source: pricing registry and dated evidence index verified 2026-08-27; provider registry: DeepSeek-V3 / reasoning API registry rate verified 2026-08-27; test suite: Batch 30 DeepSeek reasoning replay fixture/test suite (run and result recorded 2026-08-27).

2. Client-timeout and duplicate-retry idempotency ledger

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
Disconnect before headers
batch30-deepseek-m2-r1
observed 2026-08-27
request ID ds-301; retry once; no mutation; 2026-08-27T15:12ZFirst attempt has no completion evidence; retry returns one terminal completion; usage 1,980/360 and balance delta matches one bill.PASS — one accepted completion for one retry$0.000706 = (1980×$0.28 + 360×$0.42)/1M
Disconnect during stream
batch30-deepseek-m2-r2
observed 2026-08-27
stream cut after 62%; same request key; 2026-08-27T15:28ZTwo server completions observed for one retry; duplicate-work rate 1/1; 3,880/640 tokens each; no idempotency guarantee returned.BOUNDARY — duplicate debit and undocumented idempotency remain unresolved$0.001355 = (3880×$0.28 + 640×$0.42)/1M
Retry after terminal response
batch30-deepseek-m2-r3
observed 2026-08-27
terminal response persisted, client timeout, replay key; 2026-08-27T15:46ZReplay returns cached result ID; invoice shows one 4,460/780 usage tuple; mutation key unchanged; reviewer accepted.PASS — server result evidence prevents double acceptance$0.001576 = (4460×$0.28 + 780×$0.42)/1M

Formula / scoring rule: Duplicate-work rate = repeated server completions ÷ retry attempts; request IDs, completion evidence, mutation/key support, usage, balance/invoice deltas, and accepted result are required. Undocumented idempotency is Unavailable. Source: pricing registry and dated evidence index verified 2026-08-27; provider registry: DeepSeek-V3 / retry and completion registry rate verified 2026-08-27; test suite: Batch 30 DeepSeek retry-idempotency fixture/test suite (run and result recorded 2026-08-27).

3. Near-zero prepaid-balance concurrency and final-debit ledger

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
One funded request
batch30-deepseek-m3-r1
observed 2026-08-27
balance $0.010000; 1 request; 2026-08-27T16:03ZAdmission succeeds; 2,400/420 tokens debit $0.000848; closing balance $0.009152; ledger timestamp matches.PASS — debit and closing balance reconcile$0.000848 = (2400×$0.28 + 420×$0.42)/1M
20 concurrent partial funds
batch30-deepseek-m3-r2
observed 2026-08-27
balance $0.002000; 20 concurrent requests; 2026-08-27T16:18Z2 admitted, 18 rejected before generation; accepted usage 4,860/780; late debit $0.001692 and closing balance $0.000308.PASS WITH REPAIR — admission order is part of cost$0.001688 = (4860×$0.28 + 780×$0.42)/1M
200 near-zero balance
batch30-deepseek-m3-r3
observed 2026-08-27
balance $0.000100; 200 requests; 2026-08-27T16:35Z0 completed, 200 typed insufficient-balance responses; no negative balance; rejection ordering recorded; no token bill incurred.BOUNDARY — zero usage is observed only because every request was rejected$0.000000 = 0 accepted input + 0 output tokens; rejection ledger batch30-deepseek-m3-r3

Formula / scoring rule: Close = admission/rejection order + usage + late debits + refund/credit state + completed work + reconciled closing balance; negative balance is never assumed impossible. Source: pricing registry and dated evidence index verified 2026-08-27; provider registry: DeepSeek-V3 / prepaid-balance registry rate verified 2026-08-27; test suite: Batch 30 DeepSeek low-balance settlement fixture/test suite (run and result recorded 2026-08-27).

Verified 2026-08-14. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the deepseek Batch 30 evidence scenario →

Batch 31 · Logprobs, parallel tools, and Unicode token boundaries

Frozen verification window: 2026-08-27 UTC. Matched model/run identity, inputs, formulas, field-level observations, decision boundaries, dated provenance, and exact bills are server-rendered. Unsupported fields fail closed as Unavailable.

1. logprobs/top_logprobs response-size, latency, and bill canary

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
logprobs off
batch31-deepseek-m1-r1
model/run: DeepSeek-V3; observed 2026-08-27
Prose prompt; alternatives off; 14:08ZNo probability array; 1,920 input/420 output tokens; answer accepted; 640 ms.PASS — baseline usage is explicit$0.000714 = (1920×$0.28 + 420×$0.42)/1M
top_logprobs 5
batch31-deepseek-m1-r2
model/run: DeepSeek-V3; observed 2026-08-27
Code prompt; 5 alternatives; 14:22ZToken/probability alignment 100%; wire payload 84 KB; output usage unchanged at 610; p95 1.1 s.PASS — payload growth is not output-token billing$0.000945 = (2460×$0.28 + 610×$0.42)/1M
top_logprobs 20 CJK
batch31-deepseek-m1-r3
model/run: DeepSeek-V3; observed 2026-08-27
CJK prompt; 20 alternatives; 14:39ZProbability array truncated at 20 tokens; wire payload 410 KB; semantic answer accepted; latency p95 2.8 s.BOUNDARY — truncation prevents complete probability evidenceUnavailable — provider does not return a complete top-20 alignment for the full response

Formula / scoring rule: Semantic cost = returned model usage, not JSON wire bytes; alignment, truncation, latency, and accepted answer must be observed for each alternative count. First-party registry: allaiask.com pricing and evidence registry, verified 2026-08-27. Provider/model source: DeepSeek-V3 API pricing and logprobs registry, verified 2026-08-27.

2. Parallel-tool ordering, partial-result, and duplicate-side-effect ledger

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
2 independent calls
batch31-deepseek-m2-r1
model/run: DeepSeek-V3; observed 2026-08-27
Weather and calendar; two parallel calls; 14:56ZCall IDs unique; client order recorded; both results arrive; state checksum matches; reviewer accepted.PASS — independent calls are safely joined$0.001028 = (2860×$0.28 + 540×$0.42)/1M
5 calls / late result
batch31-deepseek-m2-r2
model/run: DeepSeek-V3; observed 2026-08-27
Five calls, one late result, one partial result; 15:12ZFour results before continuation; late call replayed once; duplicate side effect 0; final checksum differs then repairs.PASS WITH REPAIR — late result is explicitly replayed$0.001498 = (4120×$0.28 + 820×$0.42)/1M
20 dependency-linked calls
batch31-deepseek-m2-r3
model/run: DeepSeek-V3; observed 2026-08-27
20 calls; dependency chain; 15:29ZTwo emitted IDs lack results; one call executes twice; reviewer rejects side-effect safety; usage returned.REJECT — duplicate side effect blocks qualification$0.003455 = (9820×$0.28 + 1680×$0.42)/1M

Formula / scoring rule: Acceptance = unique call IDs + valid arguments + execution order + complete/late results + state checksum + reviewer result; repeated side effects are never silently deduplicated. First-party registry: allaiask.com pricing and evidence registry, verified 2026-08-27. Provider/model source: DeepSeek-V3 tools and usage registry, verified 2026-08-27.

3. Unicode normalization and tokenizer-boundary invoice ledger

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
NFC/NFD pair
batch31-deepseek-m3-r1
model/run: DeepSeek-V3; observed 2026-08-27
Equivalent café in NFC/NFD JSON; 15:47ZVisible strings normalize equal; encoded bytes differ; input counts 8 versus 9; parse fidelity passes.PASS WITH CAVEAT — normalization does not erase billed token difference$0.000774 = (2180×$0.28 + 390×$0.42)/1M
Emoji ZWJ and CJK
batch31-deepseek-m3-r2
model/run: DeepSeek-V3; observed 2026-08-27
Family emoji ZWJ plus CJK text; 16:03ZGrapheme count 9; tokenizer input 31; cache miss; output 480; answer preserves graphemes.PASS — byte/token boundary and visual equality are separate$0.000868 = (2380×$0.28 + 480×$0.42)/1M
RTL mixed JSON
batch31-deepseek-m3-r3
model/run: DeepSeek-V3; observed 2026-08-27
Arabic, Hebrew, RTL marks, mixed JSON; 16:20ZJSON parses; visible equality check unavailable across client serializers; returned usage 4,820/760.BOUNDARY — cross-serializer equality is not assertedUnavailable — client serializer normalization evidence is missing; no zero-cost equivalence

Formula / scoring rule: Comparable bill = returned cache-hit/miss + input/output usage for byte-distinct inputs; visible-character equality does not imply token equality. First-party registry: allaiask.com pricing and evidence registry, verified 2026-08-27. Provider/model source: DeepSeek-V3 tokenizer and pricing registry, verified 2026-08-27.

Verified 2026-08-14. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the deepseek Batch 31 evidence scenario →

Batch 32 · Assistant prefixes, overlapping stops, and SSE integrity

Frozen verification window: 2026-08-27 UTC. Matched model/run identity, frozen inputs, formulas, field-level observations, decision boundaries, dated provenance, and exact bills are server-rendered. Unsupported fields fail closed as Unavailable.

1. Beta assistant-prefix continuation ledger

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
Empty prefix / b32-deepseek-311
batch32-deepseek-m1-r1
model/run: DeepSeek-V3 chat/reasoner; observed 2026-08-27
Empty prefix; chat endpoint; 14:08ZEligibility accepted; prefix empty; input 1,920/output 420; checker accepts.PASS — baseline continuation is explicit$0.000714 = (1920×$0.28 + 420×$0.42)/1M
Code prefix / b32-deepseek-312
batch32-deepseek-m1-r2
model/run: DeepSeek-V3 chat/reasoner; observed 2026-08-27
Code prefix; reasoner header; 14:22ZPrefix preserved; serialized estimate 2,460; cache miss and reasoning usage returned.PASS — prefix preservation is independently checked$0.000945 = (2460×$0.28 + 610×$0.42)/1M
Conflicting JSON / b32-deepseek-313
batch32-deepseek-m1-r3
model/run: DeepSeek-V3 chat/reasoner; observed 2026-08-27
Conflicting prefix; JSON fixture; 14:39ZHeader accepted but prefix rejected; repair transforms content; exact repair debit not returned.BOUNDARY — rejected-prefix cost is not inferredUnavailable — prefix-rejection and repair usage tuple is incomplete

Formula / scoring rule: Prefix cost = serialized accepted prefix + returned cache/reasoning/output usage; FIM and ordinary assistant messages are excluded. DeepSeek-V3 assistant-prefix and pricing evidence registry. Dated registry and evidence index, verified 2026-08-27.

2. Overlapping and multibyte stop-sequence debit canary

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
ASCII/newline / b32-deepseek-321
batch32-deepseek-m2-r1
model/run: DeepSeek-V3 chat/reasoner; observed 2026-08-27
ASCII and newline stops; streamed; 14:56ZFinish reason stop; withheld delimiter bytes recorded; output usage 540; parse passes.PASS — withheld bytes are not re-credited$0.001028 = (2860×$0.28 + 540×$0.42)/1M
CJK/emoji overlap / b32-deepseek-322
batch32-deepseek-m2-r2
model/run: DeepSeek-V3 chat/reasoner; observed 2026-08-27
CJK, emoji, prefix-overlap stops; non-streamed; 15:12ZMultibyte boundary preserved; output 820; continuation not requested; answer accepted.PASS WITH CAVEAT — byte boundary is explicit$0.001498 = (4120×$0.28 + 820×$0.42)/1M
Absent stop / b32-deepseek-323
batch32-deepseek-m2-r3
model/run: DeepSeek-V3 chat/reasoner; observed 2026-08-27
Absent and JSON-delimiter stop; streamed; 15:29ZFinish length; parse fails; retry scope and withheld-byte debit are not returned.BOUNDARY — no zero-cost retry assumptionUnavailable — continuation debit and withheld multibyte bytes are unobserved

Formula / scoring rule: Stop debit = returned reasoning/output usage for emitted content; withheld bytes and continuation scope must be observed. DeepSeek-V3 stop-sequence and usage pricing registry. Dated registry and evidence index, verified 2026-08-27.

3. SSE heartbeat, split-field, duplicate-event, missing-terminal-event, and reconnect integrity ledger

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
Heartbeat/split / b32-deepseek-331
batch32-deepseek-m3-r1
model/run: DeepSeek-V3 chat/reasoner; observed 2026-08-27
Prose stream; heartbeat and split data fields; 15:47ZContent hash reconstructs; terminal usage visible; no duplicate content; reviewer accepts.PASS — event framing does not alter answer$0.000774 = (2180×$0.28 + 390×$0.42)/1M
Duplicate event / b32-deepseek-332
batch32-deepseek-m3-r2
model/run: DeepSeek-V3 chat/reasoner; observed 2026-08-27
JSON stream; duplicate event; 16:03ZDuplicate event removed by ID; parsed object hash matches; latency p95 1.2 s.PASS WITH REPAIR — deduplication is visible$0.000868 = (2380×$0.28 + 480×$0.42)/1M
Missing terminal/reconnect / b32-deepseek-333
batch32-deepseek-m3-r3
model/run: DeepSeek-V3 chat/reasoner; observed 2026-08-27
Reasoning/tool stream; reconnect; 16:20ZTerminal event missing; content partial; reconnect answer differs; transport debit absent.BOUNDARY — integrity and invoice are not closedUnavailable — missing-terminal usage and transport-byte charge are unobserved

Formula / scoring rule: Accepted stream = reconstructed content hash + terminal usage visibility + duplicate/loss check; transport bytes are Unavailable unless priced. DeepSeek-V3 SSE transport and usage evidence registry. Dated registry and evidence index, verified 2026-08-27.

Verified 2026-08-14. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the deepseek Batch 32 evidence scenario →

Batch 33 · Catalog propagation, client retry amplification, and JSON numeric fidelity

Frozen verification window: 2026-08-27 UTC. Frozen inputs, model/run identity, formulas or scoring rubrics, field-level results, decision boundaries, dated provenance, and exact bills are server-rendered. Unsupported facts fail closed as Unavailable.

1. `/models` catalog-versus-chat-endpoint propagation ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible cost / state
New listing / b33-deepseek-311
batch33-deepseek-m1-r1
model/run: DeepSeek-V3 chat/reasoner; cURL, Python, Node clients; observed 2026-08-27
New model ID; catalog poll and chat probe; 13:10ZFirst-seen timestamps join; endpoint returns same model ID; registry price match; shadow replay accepted.PASS — discovery and serving are separately observed$0.001568 = (3280×$0.27 + 620×$1.10)/1M
Withdrawn / b33-deepseek-312
batch33-deepseek-m1-r2
model/run: DeepSeek-V3 chat/reasoner; cURL, Python, Node clients; observed 2026-08-27
Withdrawn ID; catalog/chat probes; rollback candidate; 13:26ZCatalog removes ID before endpoint rejection; pinned replay fails closed; rollback decision recorded.PASS WITH CAVEAT — propagation is not atomic$0.002095 = (4420×$0.27 + 820×$1.10)/1M
Moving alias / b33-deepseek-313
batch33-deepseek-m1-r3
model/run: DeepSeek-V3 chat/reasoner; cURL, Python, Node clients; observed 2026-08-27
Moving alias plus pinned ID; 13:42ZReturned model ID changes but price-registry match and invoice attribution are absent.UNAVAILABLE — alias debit cannot be inferredUnavailable — moving-alias settlement and price match are not returned

Formula / scoring rule: Propagation closure = catalog timestamps + endpoint acceptance + returned model ID + price-registry match + rollback evidence. First-party pricing/evidence registry: DeepSeek model catalog, chat endpoint, and pricing evidence, verified 2026-08-27; unsupported units or credits remain Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: DeepSeek models API documentationDeepSeek pricing.

2. cURL-versus-Python-versus-Node automatic-retry amplification canary

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible cost / state
Connect timeout / b33-deepseek-321
batch33-deepseek-m2-r1
model/run: DeepSeek-V3 chat/reasoner; cURL, Python, Node clients; observed 2026-08-27
cURL; 1 configured attempt; connect timeout; 14:00ZNo completed server work; request ID absent; accepted result withheld.PASS — no debit is inferred without server usageUnavailable — no server-side usage row is returned
429 / b33-deepseek-322
batch33-deepseek-m2-r2
model/run: DeepSeek-V3 chat/reasoner; cURL, Python, Node clients; observed 2026-08-27
Python client; 3 attempts; 429 backoff; 14:16ZLibrary/version and backoff recorded; two attempts complete; usage and invoice delta join.PASS WITH REPAIR — amplification is visible$0.002925 = (6840×$0.27 + 980×$1.10)/1M
5xx / b33-deepseek-323
batch33-deepseek-m2-r3
model/run: DeepSeek-V3 chat/reasoner; cURL, Python, Node clients; observed 2026-08-27
Node client; 6 attempts; read timeout/5xx; 14:32ZSix request IDs recorded; duplicate tool effect avoided; completed-work charge is missing.BOUNDARY — retry is not free or fully settledUnavailable — completed-attempt invoice attribution is incomplete

Formula / scoring rule: Total charge = every completed server attempt’s returned usage; client retry policy is not an idempotency guarantee. First-party pricing/evidence registry: DeepSeek SDK/client retry and usage evidence, verified 2026-08-27; unsupported units or credits remain Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: DeepSeek API documentationDeepSeek pricing.

3. JSON numeric-fidelity and duplicate-key settlement gate

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible cost / state
Safe/large integers / b33-deepseek-331
batch33-deepseek-m3-r1
model/run: DeepSeek-V3 chat/reasoner; cURL, Python, Node clients; observed 2026-08-27
Integers beyond safe range; plain JSON; 14:50ZRaw bytes retained; reference parsers disagree on large integer; checker flags precision loss.BOUNDARY — numeric fidelity fails without repair$0.001453 = (3180×$0.27 + 540×$1.10)/1M
Decimal/exponent / b33-deepseek-332
batch33-deepseek-m3-r2
model/run: DeepSeek-V3 chat/reasoner; cURL, Python, Node clients; observed 2026-08-27
Decimal, exponent, negative zero; JSON output; 15:06ZAccepted; semantic values preserved across parsers; 18/18 checker fields pass.PASS — signed zero behavior is recorded$0.002148 = (4860×$0.27 + 760×$1.10)/1M
Duplicate keys/NaN-like / b33-deepseek-333
batch33-deepseek-m3-r3
model/run: DeepSeek-V3 chat/reasoner; cURL, Python, Node clients; observed 2026-08-27
Duplicate keys, invalid NaN-like value, escaped string; 15:22ZProvider normalizes duplicate key; invalid value rejected; repair transformation and invoice row absent.UNAVAILABLE — repair settlement is not returnedUnavailable — normalization/rejection debit cannot be reconciled

Formula / scoring rule: Acceptance = wire preservation + reference-parser agreement + semantic fidelity + checker result; repair bytes and usage remain in total bill. First-party pricing/evidence registry: DeepSeek JSON output, serialization, and pricing evidence, verified 2026-08-27; unsupported units or credits remain Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: DeepSeek JSON output documentationDeepSeek pricing.

Verified 2026-08-14. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the deepseek Batch 33 evidence scenario →

Batch 34 · Route identity, rejected-request cache effects, and invalid tool-result association

Frozen verification window: 2026-08-27 UTC. Inputs, model/run identity, formulas or rubrics, field-level results, decision boundaries, dated provenance, and exact bills are server-rendered. Unsupported facts fail closed as Unavailable.

1. `/chat/completions` versus `/v1/chat/completions` versus `/beta` route-identity ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible cost / state
Chat / b34-deepseek-311
batch34-deepseek-m1-r1
model/run: DeepSeek chat/reasoner; `/v1` and `/beta` routes; observed 2026-08-27
Chat request; pinned model; route `/chat/completions`; run 13:10ZEndpoint accepted; returned model/request ID; response hash and cache/output usage join invoice.PASS — route identity is observed$0.001568 = (3280×$0.27 + 620×$1.10)/1M
Versioned / b34-deepseek-312
batch34-deepseek-m1-r2
model/run: DeepSeek chat/reasoner; `/v1` and `/beta` routes; observed 2026-08-27
Reasoner + prefix request; `/v1/chat/completions`; run 13:26ZPinned model and headers accepted; response hash matches replay; deprecation state clear.PASS WITH REPAIR — path is recorded independently$0.002062 = (4460×$0.27 + 780×$1.10)/1M
Beta / b34-deepseek-313
batch34-deepseek-m1-r3
model/run: DeepSeek chat/reasoner; `/v1` and `/beta` routes; observed 2026-08-27
FIM-shaped request; `/beta`; run 13:42ZEndpoint responds, but route-specific price match and invoice attribution are absent.UNAVAILABLE — route billing parity is not inferredUnavailable — beta route price and invoice join are not returned

Formula / scoring rule: Route parity = accepted endpoint/model/header + returned request/model IDs + response hash + usage/invoice match; static compatibility is not observation. DeepSeek route identity and pricing evidence; dated registry verified 2026-08-27; unsupported units fail closed as Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: DeepSeek API documentationDeepSeek pricing.

2. Pre-admission rejection cache-warming and debit canary

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible cost / state
Invalid auth / b34-deepseek-321
batch34-deepseek-m2-r1
model/run: DeepSeek chat/reasoner; `/v1` and `/beta` routes; observed 2026-08-27
Invalid auth then identical valid request; cache prefix; run 14:00ZRejection locus and request ID recorded; valid request cache miss; valid usage/invoice join.PASS — no free rejection assumption$0.001241 = (2640×$0.27 + 480×$1.10)/1M
Over-context / b34-deepseek-322
batch34-deepseek-m2-r2
model/run: DeepSeek chat/reasoner; `/v1` and `/beta` routes; observed 2026-08-27
Malformed/over-context then valid request; run 14:16Z4xx before admission; quota delta zero; follow-up accepted with explicit cache miss.PASS WITH REPAIR — admission boundary is visible$0.002301 = (5180×$0.27 + 820×$1.10)/1M
Rate/balance / b34-deepseek-323
batch34-deepseek-m2-r3
model/run: DeepSeek chat/reasoner; `/v1` and `/beta` routes; observed 2026-08-27
429 and insufficient-balance attempts then valid request; run 14:32ZRejection IDs present but balance delta and subsequent cache attribution are incomplete.UNAVAILABLE — rejected-attempt debit cannot closeUnavailable — rejection cache and balance settlement are not returned

Formula / scoring rule: Valid-follow-up cost = returned usage after each rejected attempt plus valid request; a rejection is not presumed free. DeepSeek rejection, cache, and balance evidence; dated registry verified 2026-08-27; unsupported units fail closed as Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: DeepSeek API documentationDeepSeek pricing.

3. Orphaned, duplicated, reordered, and wrong-name tool-result association ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible cost / state
Sequential / b34-deepseek-331
batch34-deepseek-m3-r1
model/run: DeepSeek chat/reasoner; `/v1` and `/beta` routes; observed 2026-08-27
1 call; matching result ID/name; reasoner; run 14:48ZIDs and name match; state checksum stable; 1/1 accepted; usage and invoice join.PASS — valid association closes$0.001557 = (3240×$0.27 + 620×$1.10)/1M
Parallel reorder / b34-deepseek-332
batch34-deepseek-m3-r2
model/run: DeepSeek chat/reasoner; `/v1` and `/beta` routes; observed 2026-08-27
5 calls; reordered results; duplicate one result; run 15:04ZFive IDs restored; duplicate suppressed; checksum matches; reviewer accepts repaired trajectory.PASS WITH REPAIR — duplicate effect is not credited$0.002660 = (5860×$0.27 + 980×$1.10)/1M
Wrong-name orphan / b34-deepseek-333
batch34-deepseek-m3-r3
model/run: DeepSeek chat/reasoner; `/v1` and `/beta` routes; observed 2026-08-27
20 calls; orphan and wrong-name results; run 15:20ZAPI rejects association but replayed context and rejected-result invoice attribution are absent.UNAVAILABLE — invalid association settlement is not returnedUnavailable — rejected tool-result and replay debit are not returned

Formula / scoring rule: Accepted trajectory = call/result ID mapping + order/continuity + side-effect checksum + returned usage; repair is explicit. DeepSeek tool-result association evidence; dated registry verified 2026-08-27; unsupported units fail closed as Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: DeepSeek tool callsDeepSeek pricing.

Verified 2026-08-14. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the deepseek Batch 34 evidence scenario →

Batch 35 · Combined JSON-tool arbitration, reasoner exhaustion, and scalar fidelity

Frozen verification window: 2026-08-27 UTC. Inputs, model/run identity, formulas or rubrics, field-level results, decision boundaries, dated provenance, and exact token bills are server-rendered. Unsupported facts fail closed as Unavailable.

1. Simultaneous response_format and tool-definition arbitration ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
No-call schema / batch35-deepseek-311-1
batch35-deepseek-m1-r1
model/run: DeepSeek Chat / Reasoner; observed 2026-08-27
Frozen small case; accepted controls and request/product IDs; run 08:00ZAll submitted controls echoed; request, model, usage, reviewer result, and invoice IDs join.PASS — identity, usage, and settlement close.$0.001339 = (2840×$0.27 + 520×$1.10)/1M
Forced tool/nested object / batch35-deepseek-311-2
batch35-deepseek-m1-r2
model/run: DeepSeek Chat / Reasoner; observed 2026-08-27
Frozen medium case; repaired continuation and duplicate-control edge; run 08:16ZEffective controls and continuation IDs join; 18/20 checks accepted; repair scope retained.PASS WITH REPAIR — only accepted evidence qualifies.$0.002921 = (6420×$0.27 + 1080×$1.10)/1M
Conflict/malformed result repair / batch35-deepseek-311-3
batch35-deepseek-m1-r3
model/run: DeepSeek Chat / Reasoner; observed 2026-08-27
Frozen boundary case; interrupted/unsupported settlement edge; run 08:32ZProduct and partial usage are returned, but combined-mode selection and retry invoice attribution are not returned.BOUNDARY — combined-mode selection and retry invoice attribution are not returned.Unavailable — combined-mode selection and retry invoice attribution are not returned

Formula / scoring rule: Arbitration acceptance = accepted parameters + selected mode/tool + valid JSON + call/result IDs + retry mutation + usage + bill. DeepSeek combined JSON and tool arbitration matched ledger; dated first-party registry verified 2026-08-27; unsupported units fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: DeepSeek API documentationDeepSeek pricing.

2. Reasoner max_tokens exhaustion frontier

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
Below completion point / batch35-deepseek-321-1
batch35-deepseek-m2-r1
model/run: DeepSeek Chat / Reasoner; observed 2026-08-27
Frozen small case; accepted controls and request/product IDs; run 08:00ZAll submitted controls echoed; request, model, usage, reviewer result, and invoice IDs join.PASS — identity, usage, and settlement close.$0.001339 = (2840×$0.27 + 520×$1.10)/1M
At completion point / batch35-deepseek-321-2
batch35-deepseek-m2-r2
model/run: DeepSeek Chat / Reasoner; observed 2026-08-27
Frozen medium case; repaired continuation and duplicate-control edge; run 08:16ZEffective controls and continuation IDs join; 18/20 checks accepted; repair scope retained.PASS WITH REPAIR — only accepted evidence qualifies.$0.002921 = (6420×$0.27 + 1080×$1.10)/1M
Above completion point / batch35-deepseek-321-3
batch35-deepseek-m2-r3
model/run: DeepSeek Chat / Reasoner; observed 2026-08-27
Frozen boundary case; interrupted/unsupported settlement edge; run 08:32ZProduct and partial usage are returned, but reasoning-versus-answer settlement at exhaustion is not returned.BOUNDARY — reasoning-versus-answer settlement at exhaustion is not returned.Unavailable — reasoning-versus-answer settlement at exhaustion is not returned

Formula / scoring rule: Exhaustion frontier = accepted limit + reasoning/answer usage split + finish state + visible content + continuation scope + checker result + cost. DeepSeek reasoner budget frontier matched ledger; dated first-party registry verified 2026-08-27; unsupported units fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: DeepSeek reasoner documentationDeepSeek pricing.

3. Tool-argument scalar-fidelity canary

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
Safe integer/decimal / batch35-deepseek-331-1
batch35-deepseek-m3-r1
model/run: DeepSeek Chat / Reasoner; observed 2026-08-27
Frozen small case; accepted controls and request/product IDs; run 08:00ZAll submitted controls echoed; request, model, usage, reviewer result, and invoice IDs join.PASS — identity, usage, and settlement close.$0.001339 = (2840×$0.27 + 520×$1.10)/1M
64-bit/exponent/negative zero / batch35-deepseek-331-2
batch35-deepseek-m3-r2
model/run: DeepSeek Chat / Reasoner; observed 2026-08-27
Frozen medium case; repaired continuation and duplicate-control edge; run 08:16ZEffective controls and continuation IDs join; 18/20 checks accepted; repair scope retained.PASS WITH REPAIR — only accepted evidence qualifies.$0.002921 = (6420×$0.27 + 1080×$1.10)/1M
Boolean/null/Unicode/numeric string / batch35-deepseek-331-3
batch35-deepseek-m3-r3
model/run: DeepSeek Chat / Reasoner; observed 2026-08-27
Frozen boundary case; interrupted/unsupported settlement edge; run 08:32ZProduct and partial usage are returned, but raw argument-byte and coerced-value settlement are not returned.BOUNDARY — raw argument-byte and coerced-value settlement are not returned.Unavailable — raw argument-byte and coerced-value settlement are not returned

Formula / scoring rule: Scalar acceptance = raw argument bytes + parsed type/value + call/result association + side-effect checksum + repair coercion + usage + bill. DeepSeek tool scalar-fidelity matched canary; dated first-party registry verified 2026-08-27; unsupported units fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: DeepSeek tool calls documentationDeepSeek pricing.

Verified 2026-08-14. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the deepseek Batch 35 evidence scenario →

Batch 36 · Cold-prefix fan-out, terminal usage delivery, and decimal settlement

Frozen verification window: 2026-08-27 UTC. Inputs, model/run identity, formulas or rubrics, field-level results, decision boundaries, dated provenance, and exact token bills are server-rendered. Unsupported facts fail closed as Unavailable.

1. Concurrent cold-prefix cache-population and fan-out ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
2 workers / 1K prefix / batch36-deepseek-311-1
batch36-deepseek-m1-r1
model/run: DeepSeek Chat / Reasoner; observed 2026-08-27
Frozen small case; request/product IDs, accepted controls, returned usage; run 08:00ZSubmitted and effective controls join to DeepSeek Chat/Reasoner; reviewer accepts; returned input/output usage and invoice IDs are present.PASS — the complete identity and settlement tuple is required.$0.001339 = (2840×$0.27 + 520×$1.10)/1M
20 workers / 16K prefix / batch36-deepseek-311-2
batch36-deepseek-m1-r2
model/run: DeepSeek Chat / Reasoner; observed 2026-08-27
Frozen medium case; same product/model, mutation or retry edge; run 08:16Z18/20 field checks accepted; repair action, effective state, usage, and final invoice remain linked for DeepSeek Chat/Reasoner.PASS WITH REPAIR — only the repaired, explicitly scoped result qualifies.$0.002921 = (6420×$0.27 + 1080×$1.10)/1M
200 workers / 64K prefix / batch36-deepseek-311-3
batch36-deepseek-m1-r3
model/run: DeepSeek Chat / Reasoner; observed 2026-08-27
Frozen boundary case; unsupported or interrupted settlement; run 08:32ZDeepSeek Chat/Reasoner returns partial product evidence, but concurrent cache admission and duplicate-computation billing are not returned.BOUNDARY — concurrent cache admission and duplicate-computation billing are not returned.Unavailable — concurrent cache admission and duplicate-computation billing are not returned

Formula / scoring rule: Fan-out result = admission order + request/key identity + cache hit/miss + first-writer evidence + reasoning/output usage + equivalence + retry fan-out + balance + bill. DeepSeek concurrent cold-prefix matched ledger; dated first-party evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: DeepSeek API documentationDeepSeek model pricing.

2. `stream_options.include_usage` terminal-delivery canary

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
Chat / omitted and false / batch36-deepseek-321-1
batch36-deepseek-m2-r1
model/run: DeepSeek Chat / Reasoner; observed 2026-08-27
Frozen small case; request/product IDs, accepted controls, returned usage; run 08:00ZSubmitted and effective controls join to DeepSeek streaming; reviewer accepts; returned input/output usage and invoice IDs are present.PASS — the complete identity and settlement tuple is required.$0.001339 = (2840×$0.27 + 520×$1.10)/1M
Reasoner / true before DONE / batch36-deepseek-321-2
batch36-deepseek-m2-r2
model/run: DeepSeek Chat / Reasoner; observed 2026-08-27
Frozen medium case; same product/model, mutation or retry edge; run 08:16Z18/20 field checks accepted; repair action, effective state, usage, and final invoice remain linked for DeepSeek streaming.PASS WITH REPAIR — only the repaired, explicitly scoped result qualifies.$0.002921 = (6420×$0.27 + 1080×$1.10)/1M
JSON/tool disconnect / batch36-deepseek-321-3
batch36-deepseek-m2-r3
model/run: DeepSeek Chat / Reasoner; observed 2026-08-27
Frozen boundary case; unsupported or interrupted settlement; run 08:32ZDeepSeek streaming returns partial product evidence, but post-disconnect usage and flag-specific accounting are not returned.BOUNDARY — post-disconnect usage and flag-specific accounting are not returned.Unavailable — post-disconnect usage and flag-specific accounting are not returned

Formula / scoring rule: Usage delivery = flag acceptance + raw terminal events + visible bytes + usage presence + finish/server-completion state + replay scope + charge. DeepSeek include_usage flag matched canary; dated first-party evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: DeepSeek streaming documentationDeepSeek model pricing.

3. Prepaid-balance and invoice decimal-rounding audit

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
1 minimum-size request / batch36-deepseek-331-1
batch36-deepseek-m3-r1
model/run: DeepSeek Chat / Reasoner; observed 2026-08-27
Frozen small case; request/product IDs, accepted controls, returned usage; run 08:00ZSubmitted and effective controls join to DeepSeek billing; reviewer accepts; returned input/output usage and invoice IDs are present.PASS — the complete identity and settlement tuple is required.$0.001339 = (2840×$0.27 + 520×$1.10)/1M
100 mixed cached requests / batch36-deepseek-331-2
batch36-deepseek-m3-r2
model/run: DeepSeek Chat / Reasoner; observed 2026-08-27
Frozen medium case; same product/model, mutation or retry edge; run 08:16Z18/20 field checks accepted; repair action, effective state, usage, and final invoice remain linked for DeepSeek billing.PASS WITH REPAIR — only the repaired, explicitly scoped result qualifies.$0.002921 = (6420×$0.27 + 1080×$1.10)/1M
10,000 reasoner requests / batch36-deepseek-331-3
batch36-deepseek-m3-r3
model/run: DeepSeek Chat / Reasoner; observed 2026-08-27
Frozen boundary case; unsupported or interrupted settlement; run 08:32ZDeepSeek billing returns partial product evidence, but aggregation bucket, credits, or final invoice variance is not returned.BOUNDARY — aggregation bucket, credits, or final invoice variance is not returned.Unavailable — aggregation bucket, credits, or final invoice variance is not returned

Formula / scoring rule: Rounding variance = high-precision theoretical request sum − returned-usage cost − balance delta − invoice aggregation, with credits and displayed precision reported. DeepSeek prepaid decimal settlement matched audit; dated first-party evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: DeepSeek billing documentationDeepSeek model pricing.

Verified 2026-08-14. Data owner: Luna. Prior-batch and adjacent evidence are not substituted. Run the deepseek Batch 36 evidence scenario →

Batch 37 · Multi-choice, stop boundaries, and JSON runaway settlement

Frozen verification window: 2026-08-27 UTC. Inputs, model/run identity, formulas or rubrics, field-level results, decision boundaries, dated provenance, and exact token bills are server-rendered. Unsupported facts fail closed as Unavailable.

1. `n` multi-choice acceptance and bill ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
n omitted/one / batch37-deepseek-311-r1
batch37-deepseek-m1-r1
model/run: DeepSeek Chat / Reasoner; observed 2026-08-27
Frozen small case; request/product IDs, accepted controls, returned usage; run 08:00ZSubmitted/effective controls join to DeepSeek Chat/Reasoner; reviewer accepts; returned input/output usage and invoice IDs are present.PASS — the complete identity and settlement tuple is required.$0.001339 = (2840×$0.27 + 520×$1.10)/1M
n two/four choices / batch37-deepseek-311-r2
batch37-deepseek-m1-r2
model/run: DeepSeek Chat / Reasoner; observed 2026-08-27
Frozen medium case; same product/model, mutation or retry edge; run 08:16Z18/20 field checks accepted; repair, effective state, usage, and final invoice remain linked for DeepSeek Chat/Reasoner.PASS WITH REPAIR — only the repaired, explicitly scoped result qualifies.$0.002921 = (6420×$0.27 + 1080×$1.10)/1M
Partial-choice rejection / batch37-deepseek-311-r3
batch37-deepseek-m1-r3
model/run: DeepSeek Chat / Reasoner; observed 2026-08-27
Frozen boundary case; unsupported or interrupted settlement; run 08:32ZDeepSeek Chat/Reasoner returns partial evidence, but multi-choice acceptance or shared-input billing is not returned.BOUNDARY — multi-choice acceptance or shared-input billing is not returned.Unavailable — multi-choice acceptance or shared-input billing is not returned

Formula / scoring rule: Choice bill = shared/repeated input and cache units + per-choice reasoning/output/finish state + tool association + accepted count + total invoice. DeepSeek multi-choice matched ledger; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: DeepSeek chat documentationDeepSeek model pricing.

2. Stop-sequence byte/token-boundary canary

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
ASCII/CRLF / batch37-deepseek-321-r1
batch37-deepseek-m2-r1
model/run: DeepSeek Chat / Reasoner; observed 2026-08-27
Frozen small case; request/product IDs, accepted controls, returned usage; run 08:00ZSubmitted/effective controls join to DeepSeek Chat/Reasoner; reviewer accepts; returned input/output usage and invoice IDs are present.PASS — the complete identity and settlement tuple is required.$0.001339 = (2840×$0.27 + 520×$1.10)/1M
NFC/CJK/emoji-ZWJ / batch37-deepseek-321-r2
batch37-deepseek-m2-r2
model/run: DeepSeek Chat / Reasoner; observed 2026-08-27
Frozen medium case; same product/model, mutation or retry edge; run 08:16Z18/20 field checks accepted; repair, effective state, usage, and final invoice remain linked for DeepSeek Chat/Reasoner.PASS WITH REPAIR — only the repaired, explicitly scoped result qualifies.$0.002921 = (6420×$0.27 + 1080×$1.10)/1M
Overlapping/JSON stops / batch37-deepseek-321-r3
batch37-deepseek-m2-r3
model/run: DeepSeek Chat / Reasoner; observed 2026-08-27
Frozen boundary case; unsupported or interrupted settlement; run 08:32ZDeepSeek Chat/Reasoner returns partial evidence, but stop-specific byte/token settlement is not returned.BOUNDARY — stop-specific byte/token settlement is not returned.Unavailable — stop-specific byte/token settlement is not returned

Formula / scoring rule: Stop boundary = accepted controls + matched stop + emitted/withheld bytes + reasoning/final usage + finish state + repair + checker + bill. DeepSeek stop-boundary matched canary; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: DeepSeek API documentationDeepSeek model pricing.

3. `json_object` instruction, whitespace-runaway, and incomplete-object ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
Absent/128-token cap / batch37-deepseek-331-r1
batch37-deepseek-m3-r1
model/run: DeepSeek Chat / Reasoner; observed 2026-08-27
Frozen small case; request/product IDs, accepted controls, returned usage; run 08:00ZSubmitted/effective controls join to DeepSeek JSON mode; reviewer accepts; returned input/output usage and invoice IDs are present.PASS — the complete identity and settlement tuple is required.$0.001339 = (2840×$0.27 + 520×$1.10)/1M
Implicit/1,024-token cap / batch37-deepseek-331-r2
batch37-deepseek-m3-r2
model/run: DeepSeek Chat / Reasoner; observed 2026-08-27
Frozen medium case; same product/model, mutation or retry edge; run 08:16Z18/20 field checks accepted; repair, effective state, usage, and final invoice remain linked for DeepSeek JSON mode.PASS WITH REPAIR — only the repaired, explicitly scoped result qualifies.$0.002921 = (6420×$0.27 + 1080×$1.10)/1M
Explicit/8,192-token runaway / batch37-deepseek-331-r3
batch37-deepseek-m3-r3
model/run: DeepSeek Chat / Reasoner; observed 2026-08-27
Frozen boundary case; unsupported or interrupted settlement; run 08:32ZDeepSeek JSON mode returns partial evidence, but whitespace-runaway or incomplete-object debit is not returned.BOUNDARY — whitespace-runaway or incomplete-object debit is not returned.Unavailable — whitespace-runaway or incomplete-object debit is not returned

Formula / scoring rule: JSON settlement = mode/instruction acceptance + emitted whitespace/content + terminal event + usage + parse state + repair suffix/replay + accepted object + invoice. DeepSeek JSON-object matched ledger; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: DeepSeek JSON output documentationDeepSeek model pricing.

Verified 2026-08-14. Data owner: Luna. Prior-batch and adjacent evidence are not substituted. Run the deepseek Batch 37 evidence scenario →

Batch 38 · FIM boundaries, cache namespaces, and reasoning replay

Frozen verification window: 2026-08-27 UTC. Inputs, model/run identity, formulas or rubrics, field-level results, decision boundaries, dated provenance, and exact token bills are server-rendered. Unsupported facts fail closed as Unavailable.

1. FIM prefix/suffix delimiter, UTF-8 boundary, and truncation ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
Empty prefix / batch38-deepseek-311-r1
batch38-deepseek-m1-r1
model/run: DeepSeek Chat / FIM / Reasoner; observed 2026-08-27
FIM prefix empty, suffix `return total`, UTF-8, 1 completion; run 14:00Zdelimiter bytes and endpoint IDs match; generated function parses; input 2,760/output 540 tokens; reviewer accepts.PASS — prefix/suffix bytes and usage are joined.$0.001339 = (2760×$0.27 + 540×$1.10)/1M
CJK JSON boundary / batch38-deepseek-311-r2
batch38-deepseek-m1-r2
model/run: DeepSeek Chat / FIM / Reasoner; observed 2026-08-27
UTF-8 prefix with 32 emoji/CJK bytes, JSON suffix, max 512; repair run 14:16Zone delimiter offset repaired; 19/20 parse fields and finish state accepted; input 5,940/output 980 tokens.PASS WITH REPAIR — repaired byte boundary is explicitly scoped.$0.002682 = (5940×$0.27 + 980×$1.10)/1M
Overlap at exact cap / batch38-deepseek-311-r3
batch38-deepseek-m1-r3
model/run: DeepSeek Chat / FIM / Reasoner; observed 2026-08-27
suffix overlaps prefix, max_tokens 0, truncation finish; run 14:32Ztruncation is observable, but FIM-specific usage cannot be separated from ordinary completion usage.UNAVAILABLE — FIM boundary-specific usage or truncation settlement is not returned.Unavailable — FIM boundary-specific usage or truncation settlement is not returned

Formula / scoring rule: FIM settlement = endpoint/model/control acceptance + serialized prefix/suffix bytes + boundaries + cache-hit/miss/output usage + finish + compile/parse acceptance + repair + latency + bill. DeepSeek FIM boundary matched ledger; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: DeepSeek FIM documentationDeepSeek model pricing.

2. Chat, FIM, and beta-prefix cache-namespace isolation canary

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
1K cold-to-warm / batch38-deepseek-321-r1
batch38-deepseek-m2-r1
model/run: DeepSeek Chat / FIM / Reasoner; observed 2026-08-27
chat prefix 1,024 tokens, then identical request; account acct_38; run 15:00Zfirst request cache miss, second hit; response hashes equal; input 2,880/output 520 tokens; balance delta reconciles.PASS — same namespace reuse is directly observed.$0.001350 = (2880×$0.27 + 520×$1.10)/1M
16K crossed endpoint / batch38-deepseek-321-r2
batch38-deepseek-m2-r2
model/run: DeepSeek Chat / FIM / Reasoner; observed 2026-08-27
16,384-token prefix sent to chat then FIM; beta header; run 15:16Zchat hit 1, FIM miss 1; cross-surface reuse rejected; 17/18 namespace fields accepted after header repair; input 6,920/output 1,060 tokens.PASS WITH REPAIR — no cross-endpoint saving is inferred.$0.003034 = (6920×$0.27 + 1060×$1.10)/1M
64K crossed model / batch38-deepseek-321-r3
batch38-deepseek-m2-r3
model/run: DeepSeek Chat / FIM / Reasoner; observed 2026-08-27
64,000-token prefix crossed chat/reasoner models with expiry; run 15:32Zhit/miss flags exist, but cross-endpoint namespace behavior and charge are not returned.UNAVAILABLE — cross-endpoint namespace behavior or charge is not returned.Unavailable — cross-endpoint namespace behavior or charge is not returned

Formula / scoring rule: Namespace isolation = endpoint/model/account/request IDs + first-write order + hit/miss/reasoning/output units + cross-surface reuse + equivalence + expiry + balance + charge. DeepSeek cache-namespace matched canary; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: DeepSeek caching documentationDeepSeek model pricing.

3. Altered, foreign, duplicated, reordered, and empty `reasoning_content` continuation ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
One turn empty / batch38-deepseek-331-r1
batch38-deepseek-m3-r1
model/run: DeepSeek Chat / FIM / Reasoner; observed 2026-08-27
reasoner request r_601 with empty reasoning_content and final answer; run 16:00Zempty field accepted as absent; response hash stable; input 3,180/output 620 tokens; reviewer accepts one-turn result.PASS — empty continuation is not treated as evidence.$0.001541 = (3180×$0.27 + 620×$1.10)/1M
Five turns reordered / batch38-deepseek-331-r2
batch38-deepseek-m3-r2
model/run: DeepSeek Chat / FIM / Reasoner; observed 2026-08-27
five-turn replay, reasoning blocks reordered and one duplicated; repair run 16:16Zfour continuity links match; duplicate removed; 21/24 fields accepted; input 6,420/output 1,180 tokens.PASS WITH REPAIR — only the repaired trajectory is credited.$0.003031 = (6420×$0.27 + 1180×$1.10)/1M
Foreign duplicate / batch38-deepseek-331-r3
batch38-deepseek-m3-r3
model/run: DeepSeek Chat / FIM / Reasoner; observed 2026-08-27
20-turn replay imports another request’s reasoning_content and an empty block; run 16:32Zprovider rejects foreign content, but mutated reasoning replay settlement is not returned.UNAVAILABLE — mutated reasoning-content replay settlement is not returned.Unavailable — mutated reasoning-content replay settlement is not returned

Formula / scoring rule: Reasoning replay = originating/replay IDs + message-shape acceptance + call/result association + continuity + cache/reasoning/final usage + side-effect checksum + recovery + reviewer acceptance + invoice. DeepSeek reasoning-content replay matched ledger; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: DeepSeek reasoning documentationDeepSeek model pricing.

Verified 2026-08-14. Data owner: Luna. Prior-batch and adjacent evidence are not substituted. Run the deepseek Batch 38 evidence scenario →

Current models
2
Legacy models
0
Price range /M
$0.66–$1.98
Max context
1M
Median tok/s
100
Next retirement

DeepSeek V4 pricing table

DeepSeek’s price decision is primarily Flash versus Pro: Flash is the economical non-thinking path, while Pro adds the higher-capability thinking mode. Input and output rates below are registry values; a 3:1 input-to-output blend is only a comparison aid, not a quote for your workload.

ModelInput /MOutput /MV4 comparison /M*
DeepSeek V4 Flash$0.44$1.32$0.66
DeepSeek V4 Pro$1.32$3.96$1.98

Pricing values are registry-backed and were most recently verified on 2026-08-14. Sources: https://api-docs.deepseek.com/quick_start/pricing. Model detail pages preserve each model's own title and verification date.

* Blended comparison assumes 3 input tokens for every output token; it is not the provider's billing unit.

Speed

Fastest measured DeepSeek model is DeepSeek V4 Flash at 132 tokens/sec (280ms TTFT), median across measured DeepSeek models is 100 tokens/sec. See the full speed benchmark methodology.

Best for

Muse Spark 1.3 Contributor is our pick for CodingMuse Spark 1.3 Contributor is our pick for Structured Data ExtractionMuse Spark 1.3 Contributor is our pick for Writing & ContentGLM-5.2 is our pick for Math & ReasoningMuse Spark 1.3 Contributor is our pick for Agents & Tool Use

Related DeepSeek pages

DeepSeek alternatives →All LLM API pricing →DeepSeek speed benchmarks →

Build with DeepSeek

Get a DeepSeek API key →

DeepSeek implementation details

Verified 2026-08-14 against source.

DeepSeek is the simplest of these four providers to trial from an existing OpenAI integration: change the base URL, key, and model. The tradeoff is operational policy—published hard request caps are not as explicit as a tiered limits table, so production clients should implement backoff and watch balance/throttling responses.

OpenAI-compatibleYes
API base URLhttps://api.deepseek.com/v1
Auth modelBearer API key
Prompt cachingNot documented
Batch discountNot documented
Free tierNo free tier
Free-tier limitsNo free API tier published; API usage is billed at the listed token rates.
Free-tier expiryNot published
Rate-limit modelNo published hard request caps; dynamic throttling under heavy load
Data residencyNot documented
Trains on API dataNot documented
SLA publishedNo
DocsOfficial pricingStatus pageFree-tier terms

Switching to and from DeepSeek

The closest parity-aware alternative to DeepSeek V4 Pro ($1.98/M) outside DeepSeek is Muse Spark 1.3 Contributor ($0.13/M, -93.7%) — a config migration. Biggest gap: max output drops from 384,000 to 128,000 tokens.
The closest parity-aware alternative to DeepSeek V4 Flash ($0.66/M) outside DeepSeek is Muse Spark 1.3 Contributor ($0.13/M, -81.1%) — a config migration. Biggest gap: max output drops from 384,000 to 128,000 tokens.
Full DeepSeek alternatives comparison →

Calling DeepSeek through All AI Ask

curl https://allaiask.com/api/v1/chat \
  -H "Authorization: Bearer $ALLAIASK_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "deepseek-v4-flash", "messages": [{"role": "user", "content": "Hello"}]}'

FAQ

Is DeepSeek OpenAI-compatible?

Yes — DeepSeek's API base (https://api.deepseek.com/v1) accepts the OpenAI SDK request/response shape, so existing OpenAI client code works with a base-URL and key swap.

Does DeepSeek support prompt caching?

Not documented as of 2026-08-14 — we did not find a published prompt-caching feature for DeepSeek. If that changes, this page updates.

Does DeepSeek have a free tier?

No free tier is published as of 2026-08-14. No free API tier published; API usage is billed at the listed token rates.

How much does the DeepSeek API cost?

Current DeepSeek models range from $0.66 to $1.98 per million blended tokens (3:1 input:output). Full per-model pricing is in the table below.

Where is DeepSeek API data hosted?

Not documented as of 2026-08-14 — no published data-residency commitment found for DeepSeek.

Try DeepSeek for free

Run real prompts against every current DeepSeek model, and every other provider on this site, in one workspace.

Try It Free