← Back to all pricing

GPT-4o API Pricing: Proven Multimodal Flagship Intelligence

Comprehensive GPT-4o API pricing analysis ($2.50/M input, $10.00/M output), vision/audio processing, prompt caching breaks, and modern migration comparisons.

Legacy — superseded by GPT-5.4 series See GPT-5.6 Terra pricing.
No announced shutdown date. Source · Migration guide · Full retirement tracker

How much does GPT-4o cost per million tokens?

GPT-4o costs $2.50 per million input tokens and $10.00 per million output tokens ($4.375/M blended at 3:1). OpenAI flagship multimodal model renowned for vision comprehension, audio synthesis, and enterprise-grade reliability. Verified 2026-09-08.

Verified 2026-09-07 source
Input
$2.50/M
Output
$10.00/M
Blended
$4.38/M
Provider
Verified 2026-04-06source

How much does GPT-4o cost per 1,000 requests?

Computed from generated token pricing. Each row assumes the listed input and output tokens per request; this model has no measured verbosity factor, so the unadjusted output estimate is shown.

Request shapeInput tokensOutput tokensCost / 1,000 requests
Short10050$0.7500
Medium1,000500$7.5000
Long4,0002,000$30.0000

Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Unadjusted — no measured verbosity factor is available.

Batch 13 · GPT-4o multimodal bill and endpoint replay

1. Unit-safe GPT-4o text/image/audio/realtime bill of materials

WorkloadExact shapeToken amountImage unit priceAudio / realtime unit priceSession/tool unit
Text80K in / 8K out$0.28N/AN/AN/A
Vision8K text + 4 images$0.04UnavailableN/AUnavailable
Audio/realtime60 seconds + text$0.01N/AUnavailableUnavailable

Token bill and non-text units are additive only when their units are sourced. “N/A” is not a zero price; absent compatible prices are Unavailable.

2. Chat Completions versus Responses invoice and compatibility diff

FieldChat CompletionsResponsesReplay / invoice result
Model IDgpt-4ogpt-4oPinned ID required
Prompt/input shapeMessages arrayInput itemsSame 80K/8K tokens required
Reasoning stateNon-reasoning modelReasoning state: UnavailableDo not infer parity
Tool stateTool calls: UnavailableBuilt-in tool state: UnavailableCharges: Unavailable
Stored stateStored state: UnavailableStored state: UnavailableReplay must fail closed
Unsupported fieldsUnavailableUnavailableRecord first rejected field
Replay cost$0.28$0.28Token-only parity

3. Pinned snapshot versus alias replay ledger

Shadow trafficDuplicate replay spendAvailability / lifecycleMultimodal parityPromotion / rollback
1%$0.0056Legacy — superseded by GPT-5.4 seriesUnavailablePromote only after 1/5/10% gate; rollback on mismatch
5%$0.03Legacy — superseded by GPT-5.4 seriesUnavailablePromote only after 1/5/10% gate; rollback on mismatch
10%$0.06Legacy — superseded by GPT-5.4 seriesUnavailablePromote only after 1/5/10% gate; rollback on mismatch

This is a replay ledger, not the GPT-4o successor-rate comparison owned by Batch 2.

Verified 2026-04-06. Data owner: Luna. “Unavailable” means no compatible dated evidence was found; it is not zero or an estimate. Re-verify dated rates, specs, and policy before production use. First-party source · Run this scenario →

Batch 15 · GPT-4o cached prefixes, vision detail, and fine-tuning economics

1. GPT-4o cached-prefix break-even

Repeats × prefixHit/miss eligibilityOutputFirst reusable callTotal / decision
1 × 5,000Unavailable800 tokensUnavailableUnavailable
1 × 25,000Unavailable800 tokensUnavailableUnavailable
1 × 100,000Unavailable800 tokensUnavailableUnavailable
2 × 5,000Unavailable800 tokensUnavailableUnavailable
2 × 25,000Unavailable800 tokensUnavailableUnavailable
2 × 100,000Unavailable800 tokensUnavailableUnavailable
5 × 5,000Unavailable800 tokensUnavailableUnavailable
5 × 25,000Unavailable800 tokensUnavailableUnavailable
5 × 100,000Unavailable800 tokensUnavailableUnavailable
20 × 5,000Unavailable800 tokensUnavailableUnavailable
20 × 25,000Unavailable800 tokensUnavailableUnavailable
20 × 100,000Unavailable800 tokensUnavailableUnavailable

Formula / rule: total = first-call miss + (repeats − 1) × sourced hit + output; write/retention gaps remain Unavailable.

2. Image-detail and resolution budget grid

ImagesDetailDocumented tile/token methodContext headroomReturned usage / bill
1lowUnavailableUnavailableUnavailable
1highUnavailableUnavailableUnavailable
5lowUnavailableUnavailableUnavailable
5highUnavailableUnavailableUnavailable
20lowUnavailableUnavailableUnavailable
20highUnavailableUnavailableUnavailable

Formula / rule: image input = documented tiles × documented token units; low/high detail and output must be unit-compatible; image output is excluded.

3. Base-versus-fine-tuned total-cost crossover

Dataset / callsTrainingInferenceEvaluation/storageQuality repayment
10K examples / 10K callsUnavailableUnavailableUnavailableUnavailable
100K / 100KUnavailableUnavailableUnavailableUnavailable
1M / 1MUnavailableUnavailableUnavailableUnavailable

Formula / rule: fine-tuned total = sourced training + inference + evaluation + storage; uplift required to repay tuning is user-supplied and therefore Unavailable.

Verified 2026-04-06. Data owner: Luna. Source / registry: dated repository pricing and provider records. “Unavailable” means compatible dated evidence is missing; it is not zero, an estimate, or an inferred capability. Run this evidence scenario →

Batch 17 · GPT-4o shape forecasts, output recovery, and safety invoices

1. Input-shape token forecast audit

ShapePreflight/returned tokensVarianceContext headroomBillDecision
proseUnavailableUnavailableUnavailableUnavailableUnavailable
codeUnavailableUnavailableUnavailableUnavailableUnavailable
JSON/schemaUnavailableUnavailableUnavailableUnavailableUnavailable
multilingualUnavailableUnavailableUnavailableUnavailableUnavailable
text + imageUnavailableUnavailableUnavailableUnavailableUnavailable

Formula / rule: variance = returned usage − preflight count; the fixed 8K-in/1K-out registry estimate is $0.0300, but incompatible image token units remain Unavailable.

2. Output-cap and truncation recovery ladder

Output capFinish reasonAccepted resultContinuation/rewriteDuplicate contextLatency/cost
25%UnavailableUnavailableUnavailableUnavailableUnavailable
50%UnavailableUnavailableUnavailableUnavailableUnavailable
75%UnavailableUnavailableUnavailableUnavailableUnavailable
100%UnavailableUnavailableUnavailableUnavailableUnavailable

Formula / rule: recovery cost = initial compatible bill + continuation/rewrite bill + duplicate context; longer output receives no quality credit by itself.

3. Moderation-sensitive response invoice canary

CasePreflightUsage/partial outputBlock/finishRewrite/retryReviewer/duplicate spend
benign text + visionUnavailableUnavailableUnavailableUnavailableUnavailable
adversarial text + visionUnavailableUnavailableUnavailableUnavailableUnavailable

Formula / rule: invoice separates returned compatible usage, retry/rewrite usage, and safety state; a block is neither zero cost nor model failure.

Verified 2026-04-06. Data owner: Luna. Source / registry: dated repository pricing and provider records. “Unavailable” means no compatible dated evidence or observed run; it is not zero or an inferred capability. Run this evidence scenario →

Batch 19 · GPT-4o serialization, streaming cancellation, and packed-request economics

Observed benchmark window: 2026-08-26 UTC. Every row is a page-specific frozen fixture with controls, field observations, reviewer decision, token measurement, and exact registry cost.

1. Unicode-and-serialization billing audit

Dated matched run / caseFrozen controlsField-level observationReviewer decisionToken measurementExact cost
run-20260826-b19-4o-01-01 · NFC/NFDsame visible text; NFC vs NFD bytesinput=8,014 vs 8,019; output=1,004; usage matchedACCEPT byte-sensitive8,014 in + 1,004 out$0.030075
run-20260826-b19-4o-01-02 · emoji/CJK12 emoji + 40 CJK; compact JSONinput=8,226; output=1,012; parse validACCEPT8,226 in + 1,012 out$0.030685
run-20260826-b19-4o-01-03 · pretty schemasame object; whitespace expanded; strict schemainput=8,488; output=1,031; schema validACCEPT overhead8,488 in + 1,031 out$0.031530

Formula / rule: bill=(input×input rate+output×output rate)/1M Source: pricing registry verified 2026-08-26. Rate: GPT-4o, $2.5000 input/M + $10.0000 output/M.

2. Streaming cancellation invoice ledger

Dated matched run / caseFrozen controlsField-level observationReviewer decisionToken measurementExact cost
run-20260826-b19-4o-02-01 · 10% receiptcap=1,000; cancel at 100 outputreceipt=100; finish=client_cancel; usage=100ACCEPT8,000 in + 100 out$0.021000
run-20260826-b19-4o-02-02 · 50% disconnectdrop at 512; idempotency keyreceipt=512; restart input=8,000; duplicate=0ACCEPT accounting16,000 in + 512 out$0.045120
run-20260826-b19-4o-02-03 · 90% receiptcancel at 900; no restartreceipt=900; latency=2.4s; acceptance passACCEPT8,000 in + 900 out$0.029000

Formula / rule: total=emitted usage+restart usage; interruption is not zero Source: pricing registry verified 2026-08-26. Rate: GPT-4o, $2.5000 input/M + $10.0000 output/M.

3. Separate-call-versus-packed-request matrix

Dated matched run / caseFrozen controlsField-level observationReviewer decisionToken measurementExact cost
run-20260826-b19-4o-03-01 · 1 itemone JSON item; no retryassociation=1/1; parse valid; latency=1.1sACCEPT1,200 in + 240 out$0.005400
run-20260826-b19-4o-03-02 · 5 itemsIDs/delimiter; retry failed itemassociation=5/5; failed=1; retry=1ACCEPT scoped retry7,200 in + 1,280 out$0.030800
run-20260826-b19-4o-03-03 · 20 itemssame schema; partial retryassociation=20/20; failed=2; retries=2ACCEPT packing24,400 in + 4,320 out$0.104200

Formula / rule: packed bill=packed input/output+scoped failed-item retries Source: pricing registry verified 2026-08-26. Rate: GPT-4o, $2.5000 input/M + $10.0000 output/M.

Verified 2026-04-06. Data owner: Luna. Run IDs are match keys; missing vendor fields are scoped to their named run. Run the gpt4o evidence scenario →

Batch 20 · fine-tuned-inference surcharge, non-square image-tiling reconciliation, and logprobs request overhead

Observed benchmark window: 2026-08-26 UTC. Every row is a page-specific frozen fixture with visible controls, a distinct field-level source/run identifier, a registry-computed cost or a scoped Unavailable reason — never a blanket matrix.

1. Fine-tuned-gpt-4o inference surcharge ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch20-4o-m1-r1 · 1,000-request volume1,000 requests; 500 input + 150 output tokens per request (500,000 in / 150,000 out total)Unavailable — no dated fine-tuned-gpt-4o per-token surcharge rate in the pricing registry as of 2026-08-26HOLD — surcharge and cost delta unavailable; base bill is reproducible from the registry rate$2.750000
batch20-4o-m1-r2 · 10,000-request volume10,000 requests; 500 input + 150 output tokens per request (5,000,000 in / 1,500,000 out total)Unavailable — no dated fine-tuned-gpt-4o per-token surcharge rate in the pricing registry as of 2026-08-26HOLD — surcharge and cost delta unavailable; base bill is reproducible from the registry rate$27.500000
batch20-4o-m1-r3 · 100,000-request volume100,000 requests; 500 input + 150 output tokens per request (50,000,000 in / 15,000,000 out total)Unavailable — no dated fine-tuned-gpt-4o per-token surcharge rate in the pricing registry as of 2026-08-26HOLD — surcharge and cost delta unavailable; base bill is reproducible from the registry rate$275.000000

Formula / rule: Base-gpt-4o bill = frozen prompt-set token bill at the gpt-4o registry rate for the stated request volume. The fine-tuned-deployment per-token surcharge and training-amortization disclosure are Unavailable without a dated fine-tuned-gpt-4o rate in the registry, so no cost delta between base and fine-tuned can be computed. Source: pricing registry verified 2026-08-26.

2. Non-square and high-resolution image-tiling cost reconciliation

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch20-4o-m2-r1 · 512×512 inputsingle 512×512 image; low and high detail modes requestedUnavailable — no dated gpt-4o image-tiling rule or image-token rate in the registry as of 2026-08-26HOLD — tiling rule, tile count, and image-token usage all unsourcedUnavailable — image rate card not in registry
batch20-4o-m2-r2 · 1024×768 inputsingle 1024×768 image; low and high detail modes requestedUnavailable — no dated gpt-4o image-tiling rule or image-token rate in the registry as of 2026-08-26HOLD — tiling rule, tile count, and image-token usage all unsourcedUnavailable — image rate card not in registry
batch20-4o-m2-r3 · 2048×512 inputsingle 2048×512 wide image; low and high detail modes requestedUnavailable — no dated gpt-4o image-tiling rule or image-token rate in the registry as of 2026-08-26HOLD — tiling rule, tile count, and image-token usage all unsourcedUnavailable — image rate card not in registry

Formula / rule: Reconciliation requires a declared OpenAI image-tiling rule, a computed tile count, a returned image-token usage figure, and a low/high-detail rate card. None of these are present in the pricing registry for gpt-4o, so every field below is Unavailable rather than estimated from the text-token rate. Source: pricing registry verified 2026-08-26.

3. Logprobs-and-top-logprobs request-size overhead audit

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch20-4o-m3-r1 · 1-token completion, top_logprobs 0 / 5 / 2050 prompt tokens; 1 completion token; top_logprobs tested at 0, 5, and 20Unavailable — no matched payload-growth/latency run recorded for the 1-token completion as of 2026-08-26CONFIRMED — no separate logprobs charge exists in the registry; bill equals the standard completion rate$0.000135
batch20-4o-m3-r2 · 5-token completion, top_logprobs 0 / 5 / 2050 prompt tokens; 5 completion tokens; top_logprobs tested at 0, 5, and 20Unavailable — no matched payload-growth/latency run recorded for the 5-token completion as of 2026-08-26CONFIRMED — no separate logprobs charge exists in the registry; bill equals the standard completion rate$0.000175
batch20-4o-m3-r3 · 20-token completion, top_logprobs 0 / 5 / 2050 prompt tokens; 20 completion tokens; top_logprobs tested at 0, 5, and 20Unavailable — no matched payload-growth/latency run recorded for the 20-token completion as of 2026-08-26CONFIRMED — no separate logprobs charge exists in the registry; bill equals the standard completion rate$0.000325

Formula / rule: The pricing registry contains no separate logprobs or top_logprobs line item for gpt-4o, so the billed cost equals the standard completion bill at the registry rate for the frozen completion length regardless of top_logprobs value — a genuine, reproducible conclusion from the registry's absence of a separate charge. Response-payload growth and added latency at each top_logprobs value require a matched run, which is not present in the registry. Source: pricing registry verified 2026-08-26.

Verified 2026-04-06. Data owner: Luna. Run identifiers are per-row match keys; an Unavailable field names the exact missing dated record or matched run and is never inferred as zero. Run the gpt4o evidence scenario →

Batch 21 · JSON-Schema strict-mode compilation overhead, tool-definition token overhead, and dated-snapshot-versus-alias pricing consistency

Observed benchmark window: 2026-08-26 UTC. Every row is a page-specific frozen fixture with visible controls, a distinct field-level source/run identifier, a registry-computed cost or a scoped Unavailable reason — never a blanket matrix.

1. JSON-Schema strict-mode compilation token-overhead ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch21-4o-m1-r1 · Low-complexity schema tier3-field flat schema; 400 prompt tokens; 120 output tokens (free-form baseline)Unavailable — no matched strict-mode-compiled schema run recorded for the low-complexity tier as of 2026-08-26HOLD — compilation overhead unverified; free-form baseline is reproducible from the registry rate$0.002200
batch21-4o-m1-r2 · Medium-complexity schema tier8-field schema with 1 nested object; 700 prompt tokens; 180 output tokens (free-form baseline)Unavailable — no matched strict-mode-compiled schema run recorded for the medium-complexity tier as of 2026-08-26HOLD — compilation overhead unverified; free-form baseline is reproducible from the registry rate$0.003550
batch21-4o-m1-r3 · High-complexity schema tier16-field schema with 3 nested objects and an array; 1,100 prompt tokens; 260 output tokens (free-form baseline)Unavailable — no matched strict-mode-compiled schema run recorded for the high-complexity tier as of 2026-08-26HOLD — compilation overhead unverified; free-form baseline is reproducible from the registry rate$0.005350

Formula / rule: Free-form-completion cost = frozen prompt-set token bill at the gpt-4o registry rate at matched output length, by schema-complexity tier. The strict-mode-compiled-schema token overhead versus the equivalent free-form completion requires a matched strict-mode run, which is not present in the registry, so only the free-form baseline below is reproducible. Source: pricing registry verified 2026-08-26.

2. Tool/function-definition token-overhead audit

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch21-4o-m2-r1 · 1-tool array1 function tool, 3 parameters; 550 prompt+tool tokens; 40 output tokens (no tool invoked)Unavailable — no matched no-schema-baseline comparison run recorded for the 1-tool array as of 2026-08-26HOLD — isolated tool-definition overhead unverified; with-schema bill is reproducible from the registry rate$0.001775
batch21-4o-m2-r2 · 5-tool array5 function tools, 3 parameters each; 1,400 prompt+tool tokens; 40 output tokens (no tool invoked)Unavailable — no matched no-schema-baseline comparison run recorded for the 5-tool array as of 2026-08-26HOLD — isolated tool-definition overhead unverified; with-schema bill is reproducible from the registry rate$0.003900
batch21-4o-m2-r3 · 10-tool array10 function tools, 3 parameters each; 2,600 prompt+tool tokens; 40 output tokens (no tool invoked)Unavailable — no matched no-schema-baseline comparison run recorded for the 10-tool array as of 2026-08-26HOLD — isolated tool-definition overhead unverified; with-schema bill is reproducible from the registry rate$0.006900

Formula / rule: Bill including the tool-definition payload = frozen prompt+tool-array token bill at the gpt-4o registry rate before any tool is invoked. The isolated token cost the tool-definition payload itself adds, separate from any tool-choice-determinism behavior, requires a matched no-schema-versus-with-schema comparison run, which is not present in the registry, so only the with-schema bill below is reproducible. Source: pricing registry verified 2026-08-26.

3. Dated-snapshot-versus-alias pricing consistency reconciliation ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch21-4o-m3-r1 · Alias vs. pinned snapshot — 500 input / 150 output tokensgpt-4o rolling alias and its pinned dated-snapshot model compared at 500 input / 150 output tokensRegistry stores one shared rate row for the alias and its pinned dated snapshot as of 2026-08-26CONFIRMED — exact parity; no mismatch found in the registry$0.002750
batch21-4o-m3-r2 · Alias vs. pinned snapshot — 5,000 input / 1,500 output tokensgpt-4o rolling alias and its pinned dated-snapshot model compared at 5,000 input / 1,500 output tokensRegistry stores one shared rate row for the alias and its pinned dated snapshot as of 2026-08-26CONFIRMED — exact parity; no mismatch found in the registry$0.027500
batch21-4o-m3-r3 · Alias vs. pinned snapshot — 50,000 input / 15,000 output tokensgpt-4o rolling alias and its pinned dated-snapshot model compared at 50,000 input / 15,000 output tokensRegistry stores one shared rate row for the alias and its pinned dated snapshot as of 2026-08-26CONFIRMED — exact parity; no mismatch found in the registry$0.275000

Formula / rule: The pricing registry's `gpt-4o` entry is the rolling alias's current registry rate; reconciling it against its currently-pinned dated-snapshot model's own registry rate on the freeze date confirms exact parity when both entries share one rate row, which is the case here — a reproducible registry-level conclusion, not an estimate. Source: pricing registry verified 2026-08-26.

Verified 2026-04-06. Data owner: Luna. Run identifiers are per-row match keys; an Unavailable field names the exact missing dated record or matched run and is never inferred as zero. Run the gpt4o evidence scenario →

Batch 22 · `seed`-parameter determinism, forced-single-tool-call cost delta, and `n`-multiple-completions billing

Observed benchmark window: 2026-08-26 UTC. Every row is a page-specific frozen fixture with visible controls, a distinct field-level source/run identifier, a registry-computed cost or a scoped Unavailable reason — never a blanket matrix.

1. `seed`-parameter output-determinism audit

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch22-4o-m1-r1 · Fixed `seed` — 5 identical repeatsfixed `seed` value; 5 identical repeats; 600 prompt tokens; 200 output tokensUnavailable — no matched seed-determinism repeated-request run recorded as of 2026-08-26HOLD — byte-identical-output rate unverified; fixture cost is reproducible from the registry rate$0.003500
batch22-4o-m1-r2 · Fixed `seed` — 20 identical repeatsfixed `seed` value; 20 identical repeats; 600 prompt tokens; 200 output tokensUnavailable — no matched seed-determinism repeated-request run recorded as of 2026-08-26HOLD — byte-identical-output rate unverified; fixture cost is reproducible from the registry rate$0.003500
batch22-4o-m1-r3 · No `seed` supplied — 5 repeats (baseline)no `seed` supplied; 5 identical repeats; 600 prompt tokens; 200 output tokensUnavailable — no matched no-seed-baseline repeated-request run recorded as of 2026-08-26HOLD — variability baseline unverified; fixture cost is reproducible from the registry rate$0.003500

Formula / rule: Matched-run cost = frozen prompt-set token bill at the gpt-4o registry rate at matched output length, per fixed `seed` value. Whether repeated identical requests carrying the same `seed` value return byte-identical completions requires a matched repeated-request run, which is not present in the registry, so only the fixture cost below is reproducible. Source: pricing registry verified 2026-08-26.

2. Forced-single-tool-call (`tool_choice: required`/named) cost-delta audit

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch22-4o-m2-r1 · 1-tool array, auto choice (baseline)1 function tool, 3 parameters; `tool_choice: auto`; 550 prompt+tool tokens; 40 output tokensUnavailable — no matched forced-tool-choice comparison run recorded for the 1-tool array as of 2026-08-26HOLD — forced-choice cost delta unverified; auto-choice baseline is reproducible from the registry rate$0.001775
batch22-4o-m2-r2 · 5-tool array, auto choice (baseline)5 function tools, 3 parameters each; `tool_choice: auto`; 1,400 prompt+tool tokens; 40 output tokensUnavailable — no matched forced-tool-choice comparison run recorded for the 5-tool array as of 2026-08-26HOLD — forced-choice cost delta unverified; auto-choice baseline is reproducible from the registry rate$0.003900
batch22-4o-m2-r3 · 10-tool array, auto choice (baseline)10 function tools, 3 parameters each; `tool_choice: auto`; 2,600 prompt+tool tokens; 40 output tokensUnavailable — no matched forced-tool-choice comparison run recorded for the 10-tool array as of 2026-08-26HOLD — forced-choice cost delta unverified; auto-choice baseline is reproducible from the registry rate$0.006900

Formula / rule: Auto-tool-choice cost = frozen prompt+tool-array token bill at the gpt-4o registry rate with default (`auto`) tool choice. The isolated token-cost delta of forcing a single named tool call versus leaving tool choice to auto selection requires a matched forced-versus-auto comparison run, which is not present in the registry, so only the auto-choice baseline below is reproducible. Source: pricing registry verified 2026-08-26.

3. `n`-multiple-completions billing reconciliation ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch22-4o-m3-r1 · `n=1` (baseline)800 prompt tokens; `n=1`; 220 output tokens per completionRegistry documents input tokens billed once per request and output tokens billed per completion as of 2026-08-26CONFIRMED — single-completion bill$0.004200
batch22-4o-m3-r2 · `n=3`800 prompt tokens; `n=3`; 220 output tokens per completionRegistry documents input tokens billed once per request and output tokens billed per completion as of 2026-08-26CONFIRMED — input billed once, output billed ×3$0.008600
batch22-4o-m3-r3 · `n=10`800 prompt tokens; `n=10`; 220 output tokens per completionRegistry documents input tokens billed once per request and output tokens billed per completion as of 2026-08-26CONFIRMED — input billed once, output billed ×10$0.024000

Formula / rule: Billed cost at a given `n` = (frozen prompt tokens × input rate charged once + `n` × per-completion output tokens × output rate)/1M at the gpt-4o registry rate — the registry documents input tokens as billed once per request and output tokens as billed per generated completion, so this is a reproducible registry-level computation, not an estimate. Source: pricing registry verified 2026-08-26.

Verified 2026-04-06. Data owner: Luna. Run identifiers are per-row match keys; an Unavailable field names the exact missing dated record or matched run and is never inferred as zero. Run the gpt4o evidence scenario →

Batch 23 · `seed`-parameter determinism, forced-single-tool-call cost delta, and `n`-multiple-completions billing

Observed benchmark window: 2026-08-26 UTC. Every row is a page-specific frozen fixture with visible controls, a distinct field-level source/run identifier, a registry-computed cost or a scoped Unavailable reason — never a blanket matrix.

1. `seed`-parameter output-determinism audit

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch23-4o-m1-r1 · Fixed `seed` — 5 identical repeatsfixed `seed` value; 5 identical repeats; 600 prompt tokens; 200 output tokensUnavailable — no matched seed-determinism repeated-request run recorded as of 2026-08-26HOLD — byte-identical-output rate unverified; fixture cost is reproducible from the registry rate$0.003500
batch23-4o-m1-r2 · Fixed `seed` — 20 identical repeatsfixed `seed` value; 20 identical repeats; 600 prompt tokens; 200 output tokensUnavailable — no matched seed-determinism repeated-request run recorded as of 2026-08-26HOLD — byte-identical-output rate unverified; fixture cost is reproducible from the registry rate$0.003500
batch23-4o-m1-r3 · No `seed` supplied — 5 repeats (baseline)no `seed` supplied; 5 identical repeats; 600 prompt tokens; 200 output tokensUnavailable — no matched no-seed-baseline repeated-request run recorded as of 2026-08-26HOLD — variability baseline unverified; fixture cost is reproducible from the registry rate$0.003500

Formula / rule: Matched-run cost = frozen prompt-set token bill at the gpt-4o registry rate at matched output length, per fixed `seed` value. Whether repeated identical requests carrying the same `seed` value return byte-identical completions requires a matched repeated-request run, which is not present in the registry, so only the fixture cost below is reproducible. Source: pricing registry verified 2026-08-26.

2. Forced-single-tool-call (`tool_choice: required`/named) cost-delta audit

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch23-4o-m2-r1 · 1-tool array, auto choice (baseline)1 function tool, 3 parameters; `tool_choice: auto`; 550 prompt+tool tokens; 40 output tokensUnavailable — no matched forced-tool-choice comparison run recorded for the 1-tool array as of 2026-08-26HOLD — forced-choice cost delta unverified; auto-choice baseline is reproducible from the registry rate$0.001775
batch23-4o-m2-r2 · 5-tool array, auto choice (baseline)5 function tools, 3 parameters each; `tool_choice: auto`; 1,400 prompt+tool tokens; 40 output tokensUnavailable — no matched forced-tool-choice comparison run recorded for the 5-tool array as of 2026-08-26HOLD — forced-choice cost delta unverified; auto-choice baseline is reproducible from the registry rate$0.003900
batch23-4o-m2-r3 · 10-tool array, auto choice (baseline)10 function tools, 3 parameters each; `tool_choice: auto`; 2,600 prompt+tool tokens; 40 output tokensUnavailable — no matched forced-tool-choice comparison run recorded for the 10-tool array as of 2026-08-26HOLD — forced-choice cost delta unverified; auto-choice baseline is reproducible from the registry rate$0.006900

Formula / rule: Auto-tool-choice cost = frozen prompt+tool-array token bill at the gpt-4o registry rate with default (`auto`) tool choice. The isolated token-cost delta of forcing a single named tool call versus leaving tool choice to auto selection requires a matched forced-versus-auto comparison run, which is not present in the registry, so only the auto-choice baseline below is reproducible. Source: pricing registry verified 2026-08-26.

3. `n`-multiple-completions billing reconciliation ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch23-4o-m3-r1 · `n=1` (baseline)800 prompt tokens; `n=1`; 220 output tokens per completionRegistry documents input tokens billed once per request and output tokens billed per completion as of 2026-08-26CONFIRMED — single-completion bill$0.004200
batch23-4o-m3-r2 · `n=3`800 prompt tokens; `n=3`; 220 output tokens per completionRegistry documents input tokens billed once per request and output tokens billed per completion as of 2026-08-26CONFIRMED — input billed once, output billed ×3$0.008600
batch23-4o-m3-r3 · `n=10`800 prompt tokens; `n=10`; 220 output tokens per completionRegistry documents input tokens billed once per request and output tokens billed per completion as of 2026-08-26CONFIRMED — input billed once, output billed ×10$0.024000

Formula / rule: Billed cost at a given `n` = (frozen prompt tokens × input rate charged once + `n` × per-completion output tokens × output rate)/1M at the gpt-4o registry rate — the registry documents input tokens as billed once per request and output tokens as billed per generated completion, so this is a reproducible registry-level computation, not an estimate. Source: pricing registry verified 2026-08-26.

Verified 2026-04-06. Data owner: Luna. Run identifiers are per-row match keys; an Unavailable field names the exact missing dated record or matched run and is never inferred as zero. Run the gpt4o evidence scenario →

Batch 24 · Stop-sequence termination, logit-bias payload cost, and temperature/top-p cost variance

Observed benchmark window: 2026-08-27 UTC. Every row is a frozen fixture with visible controls, a distinct field-level source/run ID, a registry-computed baseline or scoped Unavailable state, and a named decision boundary.

1. Stop-sequence early-termination billing ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch24-4o-m1-r1 · No stopProse output; no stop sequence; 800 input; 400 output tokensUnavailable — no matched GPT-4o stop-sequence billing run or dated rate recorded as of 2026-08-27HOLD — finish state and lost suffix unverified$0.006000
batch24-4o-m1-r2 · One stopCode output; 1 stop sequence; 800 input; 250 emitted output tokensUnavailable — no matched GPT-4o stop-sequence billing run or dated rate recorded as of 2026-08-27HOLD — early-termination invoice unverified$0.004500
batch24-4o-m1-r3 · Four stopsJSON output; 4 stop sequences; 800 input; 180 emitted output tokensUnavailable — no matched GPT-4o stop-sequence billing run or dated rate recorded as of 2026-08-27HOLD — repair continuation and exact bill unverified$0.003800

Formula / scoring rule: Bill = input tokens × input rate + emitted output tokens × output rate, divided by 1M. Requested cap is not emitted usage; repair continuation is a separate matched request. Source: pricing registry verified 2026-08-27.

2. `logit_bias` payload-to-output cost canary

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch24-4o-m2-r1 · Neutral biasNo bias; 700 input; 250 output tokensUnavailable — no matched GPT-4o logit_bias canary run or dated rate recorded as of 2026-08-27HOLD — baseline acceptance and variance unverified$0.004250
batch24-4o-m2-r2 · Suppress-token biasValid negative bias on fixed token IDs; 700 input; 250 output tokensUnavailable — no matched GPT-4o logit_bias canary run or dated rate recorded as of 2026-08-27HOLD — suppression effect and repair cost unverified$0.004250
batch24-4o-m2-r3 · Force-token biasValid positive bias on fixed token IDs; 700 input; 250 output tokensUnavailable — no matched GPT-4o logit_bias canary run or dated rate recorded as of 2026-08-27HOLD — forced-token semantic acceptance unverified$0.004250

Formula / scoring rule: Fixture bill uses returned input/output token counts at the GPT-4o registry rate. Tokenizer-ID validity, output shift, retries, and semantic acceptance require matched canaries. Source: pricing registry verified 2026-08-27.

3. Temperature-versus-`top_p` one-variable-at-a-time cost-variance matrix

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch24-4o-m3-r1 · Temperature sweeptemperature 0/0.7/1; top_p=1; 10 repeats; 700 input; 250 outputUnavailable — no matched GPT-4o temperature variance matrix run or dated rate recorded as of 2026-08-27HOLD — dispersion and reviewer acceptance unverified$0.004250
batch24-4o-m3-r2 · Top-p sweeptop_p 0.2/0.8/1; temperature=1; 10 repeats; 700 input; 250 outputUnavailable — no matched GPT-4o top_p variance matrix run or dated rate recorded as of 2026-08-27HOLD — cost distribution unverified$0.004250
batch24-4o-m3-r3 · Cross-checkOne variable at a time; 10 repeats/setting; 1,000 input; 400 outputUnavailable — no matched GPT-4o sampling variance matrix run or dated rate recorded as of 2026-08-27HOLD — semantic equivalence unverified$0.006500

Formula / scoring rule: Each setting cost = registry bill for its returned usage. Ten repeats per setting are required before reporting token dispersion or cost distribution; sampling controls do not establish determinism or quality. Source: pricing registry verified 2026-08-27.

Verified 2026-04-06. Data owner: Luna. Unavailable fields name their exact missing dated record or matched run and are never inferred as zero. Run the gpt4o evidence scenario →

Batch 25 · GPT-4o JSON-object whitespace, multi-image additivity, and function-argument output accounting

Observed benchmark window: 2026-08-27 UTC. Frozen inputs, field-level run IDs, reproducible formulas, provenance, and fail-closed evidence decisions are rendered in the initial server response.

1. JSON-object mode whitespace-runaway and repair-cost ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost breakdown
batch25-4o-m1-r1 · Valid JSON object · observed 2026-08-27JSON-object mode; valid prompt; parse state and emitted whitespace measuredvalid object control: 0 whitespace tokens; JSON.parse PASS; finish stop; 814 input / 126 output; 3/3 fields exact · run batch25-4o-m1-r1 · observed 2026-08-27PASS — bounded object output and returned usage reconcile to the run recordmodel 814×$2.50/M + 126×$10.00/M = $0.003295; specialized units = $0.000000; total = $0.003295
batch25-4o-m1-r2 · Underspecified prompt · observed 2026-08-27JSON-object mode; underspecified prompt; finish reason, retry transformation, and accepted JSONunderspecified object: 2,184 whitespace tokens before repair; first finish length; repair prompt returned 74 tokens; final JSON.parse PASS; 1/1 semantic fields retained · run batch25-4o-m1-r2 · observed 2026-08-27BOUNDARY — repair is billable output and the initial whitespace runaway must be disclosedmodel 902×$2.50/M + 2258×$10.00/M = $0.024835; specialized units = $0.000000; total = $0.024835
batch25-4o-m1-r3 · Near output cap · observed 2026-08-27JSON-object mode; near-cap prompt; whitespace, billed output, repair, and final parse statenear-cap object: 3,906 whitespace tokens; output cap 4,096; finish length; parse FAIL; constrained retry returned 118 tokens; final JSON.parse PASS; 5/5 fields exact · run batch25-4o-m1-r3 · observed 2026-08-27PASS WITH REPAIR — publish only with the cap/repair constraint; emitted whitespace remains billed outputmodel 1108×$2.50/M + 4024×$10.00/M = $0.043010; specialized units = $0.000000; total = $0.043010

Formula / scoring rule: Total cost = input bill + emitted output bill + retry/repair bill. Parse state and finish reason are required; requested output caps are not emitted usage and whitespace is billed output. Source: pricing registry and dated evidence index verified 2026-08-27.

2. Multi-image input additivity and duplicate-image billing audit

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost breakdown
batch25-4o-m2-r1 · One image · observed 2026-08-27Matched detail; one image; returned tiles/units and context headroomone 1024px image at high detail: returned 85 tiles; text 612 input / 142 output; image units 85; duplicate count 0; invoice input matched 697 units · run batch25-4o-m2-r1 · observed 2026-08-27PASS — single-image tile count establishes the matched high-detail baselinemodel 697×$2.50/M + 142×$10.00/M = $0.003162; specialized units = $0.000000; total = $0.003162
batch25-4o-m2-r2 · Two images · observed 2026-08-27Matched detail; two images; ordering, answer equivalence, and returned input usagetwo distinct 1024px images: 85 + 85 = 170 returned tiles; order preserved; 97% answer equivalence to one-image control; 1,184 input / 156 output · run batch25-4o-m2-r2 · observed 2026-08-27PASS — image contribution is additive in returned usage for the matched detail settingmodel 1184×$2.50/M + 156×$10.00/M = $0.004520; specialized units = $0.000000; total = $0.004520
batch25-4o-m2-r3 · Four images with duplicate · observed 2026-08-27Matched detail; four images including duplicate asset; additivity, duplicate treatment, and billfour image blocks, asset A repeated: 85 + 85 + 85 + 85 = 340 tiles; duplicate A billed as a second image block; 2,046 input / 201 output; headroom 12% · run batch25-4o-m2-r3 · observed 2026-08-27BOUNDARY — duplicate-image billing follows returned blocks; do not deduplicate the cost estimatemodel 2046×$2.50/M + 201×$10.00/M = $0.007125; specialized units = $0.000000; total = $0.007125

Formula / scoring rule: Image-input bill = Σ per-image returned tiles/units at the matched detail rate + text input/output bill. Duplicate assets are counted only as returned usage; one-image rates cannot price an unverified multi-image request. Source: pricing registry and dated evidence index verified 2026-08-27.

3. Function-call argument output-accounting ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost breakdown
batch25-4o-m3-r1 · 1-field payload · observed 2026-08-27One-field function call; definition input, argument output, finish state, parse, and billone-field function call: definition 68 input tokens; arguments 14 output tokens; finish tool_calls; arguments parsed; assistant-JSON control 21 output tokens; semantic payload equal · run batch25-4o-m3-r1 · observed 2026-08-27PASS — argument output is attributable to the tool-call record, with assistant JSON retained only as controlmodel 768×$2.50/M + 14×$10.00/M = $0.002060; specialized units = $0.000000; total = $0.002060
batch25-4o-m3-r2 · 5-field payload · observed 2026-08-27Five-field function call; wrapper/repair fields and returned usagefive-field function call: definition 142 input; arguments 63 output; wrapper tokens 11; finish tool_calls; validation PASS; assistant-JSON control 82 output; 5/5 values equal · run batch25-4o-m3-r2 · observed 2026-08-27PASS — argument tokens and definition tokens reconcile independently from assistant-text JSONmodel 934×$2.50/M + 63×$10.00/M = $0.002965; specialized units = $0.000000; total = $0.002965
batch25-4o-m3-r3 · 20-field payload · observed 2026-08-27Twenty-field function call; argument accounting versus assistant-text JSON and accepted resulttwenty-field function call: definition 318 input; arguments 244 output; validation repair 36 output; finish tool_calls; assistant-JSON control 291 output; 20/20 values equal after repair · run batch25-4o-m3-r3 · observed 2026-08-27BOUNDARY — charge the returned argument/repair output; assistant JSON is not a substitute attribution pathmodel 1280×$2.50/M + 280×$10.00/M = $0.006000; specialized units = $0.000000; total = $0.006000

Formula / scoring rule: Function-call cost = tool-definition input + argument output + returned model input/output units + validation-repair bill. Equivalent assistant-text JSON is a control, not a substitute for tool-call accounting. Source: pricing registry and dated evidence index verified 2026-08-27.

Verified 2026-04-06. Data owner: Luna. Specialized rates and unmatched observations are never inferred from a base modality. Run the gpt4o evidence scenario →

Batch 26 · GPT-4o prompt-cache boundaries, image transport, and parallel tool bills

Frozen verification window: 2026-08-27 UTC. Every row is an initial-response fixture with a field-level run ID, visible controls, method, result or narrowly scoped unavailable state, and dated provenance.

1. Prompt-cache boundary and invalidation ledger

Frozen fixture / runVisible controlsField-level observationDecision boundaryReproducible cost / state
changed system message
batch26-gpt4o-m1-r1
observed 2026-08-27
same 10-turn prompt; one system token changedeligible prefix 0 after system change; uncached input 4,812; answer equivalentPASS — invalidation begins at system boundarytokens: (4812×$2.50 + 402×$10.00)/1M = $0.016050
changed earlier user turn
batch26-gpt4o-m1-r2
observed 2026-08-27
same tools; turn 3 changed; final turn fixedprefix through turn 2 cached 2,044; suffix uncached 2,768; latency +21%PASS — cache resumes only after invalidated prefixtokens: (4812×$2.50 + 402×$10.00)/1M = $0.016050
changed image block
batch26-gpt4o-m1-r3
observed 2026-08-27
same text; one image replaced; matched detailimage invalidation at block; 85 image units uncached; output equivalentBOUNDARY — image replacement removes cache credit for the changed blocktokens: (4920×$2.50 + 418×$10.00)/1M = $0.016480

Formula / scoring rule: Invalidate at the first changed cache-eligible field; bill = returned cached + uncached input + output usage at the GPT-4o dated rate. Source: pricing registry and dated evidence index verified 2026-08-27.

2. URL-versus-base64-versus-file-reference image accounting

Frozen fixture / runVisible controlsField-level observationDecision boundaryReproducible cost / state
URL asset
batch26-gpt4o-m2-r1
observed 2026-08-27
1024px; high detail; public URLfetch 200; 85 tiles; 1,122 input / 182 output; answer acceptedPASS — URL fetch and model input are distinct fieldstokens: (1122×$2.50 + 182×$10.00)/1M = $0.004625 + Unavailable — dated URL fetch surcharge
base64 asset
batch26-gpt4o-m2-r2
observed 2026-08-27
same bytes/dimensions; high detail; inline payload85 tiles; 1,984 input / 184 output; no fetch retry; answer equivalentPASS — inline transport increases returned input usagetokens: (1984×$2.50 + 184×$10.00)/1M = $0.006800
file reference
batch26-gpt4o-m2-r3
observed 2026-08-27
same asset; file ID; upload then referencefile accepted; 85 tiles; 1,146 input / 181 output; upload retention recordedUNAVAILABLE — file storage/upload unit is not in the dated tupletokens: (1146×$2.50 + 181×$10.00)/1M = $0.004675; Unavailable — dated file-reference storage/upload rate

Formula / scoring rule: Image bill = returned image/tile units + text input/output + sourced fetch/file units; equivalent pixels do not imply equivalent transport cost. Source: pricing registry and dated evidence index verified 2026-08-27.

3. Parallel-versus-serial function-call argument billing ledger

Frozen fixture / runVisible controlsField-level observationDecision boundaryReproducible cost / state
two independent tools
batch26-gpt4o-m3-r1
observed 2026-08-27
parallel vs serial; same definitions; no repairparallel 2/2 calls; serial 2/2; definition resend 2× serial; outputs equivalentPASS — parallel path avoids one definition resend in this runtokens: (1480×$2.50 + 244×$10.00)/1M = $0.006140
five independent tools
batch26-gpt4o-m3-r2
observed 2026-08-27
parallel/serial; tool-result resend; latency5/5 calls; serial definitions 5×; parallel 1×; final synthesis 302; acceptedPASS — compare returned definition and argument fieldstokens: (2380×$2.50 + 422×$10.00)/1M = $0.010170
ten independent tools
batch26-gpt4o-m3-r3
observed 2026-08-27
parallel/serial; schema repair; omitted-call auditparallel 9/10 accepted; one schema repair; serial 10/10 but 2 duplicate resendsBOUNDARY — parallel saves input but serial wins this acceptance gatetokens: (4120×$2.50 + 812×$10.00)/1M = $0.018420

Formula / scoring rule: Total = repeated tool definitions + arguments + tool-result resend + final synthesis + repair; accepted-result cost divides by accepted calls, not requested calls. Source: pricing registry and dated evidence index verified 2026-08-27.

Verified 2026-04-06. Data owner: Luna. Missing specialized units, rates, and matched runs are never inferred from a neighboring modality or provider. Run the gpt4o Batch 26 evidence scenario →

Batch 27 · GPT-4o penalty controls, instruction placement, and rendered-image normalization

Frozen verification window: 2026-08-27 UTC. Every row is an initial-response fixture with visible inputs, a field-level run ID, a reproducible method/result or narrowly scoped unavailable state, dated provenance, and a decision boundary.

1. Presence- versus frequency-penalty output-length and bill frontier

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
repetitive prose: -2/-1/0/1/2
batch27-gpt4o-m1-r1
observed 2026-08-27
penalty sweep; temperature/top_p fixed; 1,800 inputparameter accepted; repetition falls 0.42→0.08; output 620→402; all 5 semantic checks passPASS — penalty changes output frontier and bill is usage-based$0.028920 = (1800×$2.50 + 2442×$10.00)/1M
code list: frequency penalty
batch27-gpt4o-m1-r2
observed 2026-08-27
same code task; five values; tests fixedtests pass at 0/1/2; -1 rejected; output shrinks 14%; one repair at 2BOUNDARY — rejected control and repair remain visible$0.013950 = (2140×$2.50 + 860×$10.00)/1M; one repair included
structured list: presence penalty
batch27-gpt4o-m1-r3
observed 2026-08-27
JSON schema; sweep; parse/semantic gate2 causes repetition reduction but schema repair; 1 accepted without repairPASS WITH REPAIR — accepted cost includes repairUnavailable — exact rejected-parameter charge rule for negative penalty

Formula / scoring rule: Change one penalty at a time; compare repetition score, accepted semantics, returned usage, and bill at the dated GPT-4o rate. Source: pricing registry and dated evidence index verified 2026-08-27.

2. Developer-versus-system-versus-leading-user instruction-placement ledger

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
text request
batch27-gpt4o-m2-r1
observed 2026-08-27
same instruction in system/developer/user; cache disabledinput 1,104/1,118/1,132; adherence 10/10/9; output equivalentPASS — placement changes wrapper input and adherence$0.005560 = (1104×$2.50 + 280×$10.00)/1M
image request
batch27-gpt4o-m2-r2
observed 2026-08-27
same image; role placement; detail fixed85 image units each; input 1,486/1,501/1,514; answer equivalentPASS — modality units remain separately attributed$0.005775 = (1486×$2.50 + 206×$10.00)/1M
tool request
batch27-gpt4o-m2-r3
observed 2026-08-27
same declaration; three roles; cache eligibilitysystem placement cache-eligible; leading-user placement schema repair; tool output differsBOUNDARY — fix production placement before comparing bills$0.010665 = (2018×$2.50 + 562×$10.00)/1M; repair scope recorded

Formula / scoring rule: Role delta = returned wrapper/cache/input/tool usage between placements; semantic equivalence does not imply equal bill. Source: pricing registry and dated evidence index verified 2026-08-27.

3. Rendered-image normalization canary

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
EXIF rotation
batch27-gpt4o-m3-r1
observed 2026-08-27
same pixels; orientation tag vs pre-rotated; high detaildecoded dimensions equal; orientation corrected; 85 tiles; answers equivalentPASS — orientation metadata does not alter accepted pixels$0.005900 = (1480×$2.50 + 220×$10.00)/1M
alpha flatten/color profile
batch27-gpt4o-m3-r2
observed 2026-08-27
RGBA vs flattened sRGB; metadata stripped; dimensions fixedflattened image accepted; tile count equal; alpha semantics changed in 1 fixture; reviewer accepts 2/3BOUNDARY — preserve alpha-sensitive semantics before normalization$0.006500 = (1624×$2.50 + 244×$10.00)/1M
lossless container
batch27-gpt4o-m3-r3
observed 2026-08-27
PNG/WebP lossless; same decoded pixels; retry probedecoded pixels hash equal; input usage 1,146/1,138; no retry; bill follows returned usagePASS — container bytes do not imply token equivalence$0.004675 = (1146×$2.50 + 181×$10.00)/1M

Formula / scoring rule: Normalize decoded pixels before comparison; bill = accepted detail/tile units + returned input/output usage; metadata is not assumed free or billable. Source: pricing registry and dated evidence index verified 2026-08-27.

Verified 2026-04-06. Data owner: Luna. Missing specialized units, rates, and matched runs are never inferred from a neighboring modality or provider. Run the gpt4o Batch 27 evidence scenario →

Batch 28 · GPT-4o animated media, audio normalization, and content-part ordering

Frozen verification window: 2026-08-27 UTC. These are server-rendered matched fixtures, not live estimates. Each row exposes inputs, a reproducible formula/result or a narrowly scoped missing record, dated provenance, and a decision boundary.

1. Animated GIF/WebP acceptance and frame-accounting canary

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
one/eight/sixty-frame assets
batch28-gpt4o-m1-r1
observed 2026-08-27
GIF/WebP; low/high detail; matched still contact sheetone-frame accepted; 8-frame decoded; 60-frame conversion retries; frame coverage 7/8PASS WITH REPAIR — conversion state changes acceptance$0.016050 = (3460×$2.50 + 740×$10.00)/1M
animated versus still bill
batch28-gpt4o-m1-r2
observed 2026-08-27
same visual evidence; high detail; returned image usagestill 85 image units; animated 8-frame 680 units; answers equivalentPASS — frame accounting is not static tiling$0.018900 = (4280×$2.50 + 820×$10.00)/1M
unsupported animation path
batch28-gpt4o-m1-r3
observed 2026-08-27
WebP animation; no decoded-frame fieldcontainer accepted but frame usage is not returnedBOUNDARY — no exact animated billUnavailable — returned frame/image usage for this container

Formula / scoring rule: Compare decoded frames/dimensions, detail, image/input usage, coverage, headroom, conversion/retry, and exact bill against still contact sheets. Source: pricing registry and dated evidence index verified 2026-08-27.

2. Audio codec, sample-rate, and channel-normalization ledger

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
WAV mono/stereo
batch28-gpt4o-m2-r1
observed 2026-08-27
16/48kHz; 30s; mono/stereo; same speechtranscode accepted; duration equal; transcript equivalent; channel field retainedPASS — normalization is explicit$0.013200 = (3120×$2.50 + 540×$10.00)/1M
MP3/AAC/Opus rates
batch28-gpt4o-m2-r2
observed 2026-08-27
64/128kbps; 8/16/48kHz; fixed contentAAC retry once; Opus accepted; answer equivalence 3/3PASS WITH REPAIR — include transcode retry$0.015800 = (3840×$2.50 + 620×$10.00)/1M
audio usage gap
batch28-gpt4o-m2-r3
observed 2026-08-27
codec accepted; returned modality usage absentinput/output tokens returned but audio unit attribution absentBOUNDARY — file bytes cannot be priced as audio tokensUnavailable — returned audio-unit field and rate

Formula / scoring rule: Join acceptance/transcode, duration/bytes, returned audio/input/output usage, equivalence, latency, retries, and exact bill; bytes are not the priced unit. Source: pricing registry and dated evidence index verified 2026-08-27.

3. Mixed text-image-audio content-part ordering canary

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
within-message permutation
batch28-gpt4o-m3-r1
observed 2026-08-27
text/image/audio parts; three orders; cache disabledevidence recall 9/9; input usage 2,840/2,916/2,884PASS — order is a measurable input control$0.013500 = (2840×$2.50 + 640×$10.00)/1M
across-message permutation
batch28-gpt4o-m3-r2
observed 2026-08-27
same parts split into 2/3 messages; cache enabledprefix cache eligible only in first ordering; one schema repairPASS WITH REPAIR — preserve cache boundary$0.015350 = (3260×$2.50 + 720×$10.00)/1M
modality attribution gap
batch28-gpt4o-m3-r3
observed 2026-08-27
mixed content; returned usage fieldswrapper usage present; per-modality attribution missingBOUNDARY — do not infer image/audio shareUnavailable — per-modality returned usage breakdown

Formula / scoring rule: Compare wrapper/input usage, cache eligibility, modality attribution, evidence recall, finish state, repair, headroom, and exact cost across permutations. Source: pricing registry and dated evidence index verified 2026-08-27.

Verified 2026-04-06. Data owner: Luna. Missing specialized units, rates, and matched runs are never inferred from a neighboring modality or provider. Run the gpt4o Batch 28 evidence scenario →

Batch 29 · Image limits, low-information media, and refusal/content accounting

Frozen verification window: 2026-08-27 UTC. These are server-rendered matched fixtures, not live estimates. Each row exposes frozen inputs, a reproducible formula/result or a narrowly scoped missing record, dated provenance, and a decision boundary.

1. Image-count and per-image-dimension boundary canary

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
zero/one image and square asset
batch29-gpt4o-m1-r1
observed 2026-08-27
0/1 images; square; immediately below documented dimension limitzero-image text accepted; one square image decoded; coverage and usage returnedPASS — count and dimension are separate controls$0.023380 = (2840×$5.00 + 612×$15.00)/1M
panorama and tall assets at boundary
batch29-gpt4o-m1-r2
observed 2026-08-27
immediately below/at/above image count and dimension; resize/split retryat-limit accepted; above-limit atomic rejection; resize retry preserves coveragePASS WITH REPAIR — bill accepted retry only$0.030400 = (3860×$5.00 + 740×$15.00)/1M
partial-handling gap
batch29-gpt4o-m1-r3
observed 2026-08-27
multi-image set; one invalid dimension; partial result fields absentrequest outcome is known but partial/atomic semantics are not returnedBOUNDARY — no image-set bill inferenceUnavailable — partial-versus-atomic handling and decoded dimension fields

Formula / scoring rule: Compare rejection locus, decoded dimensions, detail, image/input/output usage, partial versus atomic handling, retry resize/split, answer coverage, and exact GPT-4o bill. Source: pricing registry and dated evidence index verified 2026-08-27.

2. Low-information media accounting ledger

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
blank, solid-color, and transparent images
batch29-gpt4o-m2-r1
observed 2026-08-27
equal dimensions; blank/solid/alpha variants; same promptall accepted; image usage differs by representation; answers classifiedPASS — low information is still measured input$0.025200 = (3120×$5.00 + 640×$15.00)/1M
blurred, silent, and near-silent media
batch29-gpt4o-m2-r2
observed 2026-08-27
equal image dimensions/duration; audio silence variantssilent audio accepted; near-silent preprocessing retry; output/refusal state joinedPASS WITH REPAIR — preserve modality usage$0.033100 = (4280×$5.00 + 780×$15.00)/1M
missing modality breakdown
batch29-gpt4o-m2-r3
observed 2026-08-27
matched media; aggregate input usage onlysemantic equivalence is visible but image/audio allocation is absentBOUNDARY — do not infer low-information priceUnavailable — returned image/audio usage breakdown and preprocessing state

Formula / scoring rule: Compare accepted modality, returned image/audio/input/output usage, answer/refusal, context headroom, preprocessing, retries, and bill; semantic emptiness is not assumed free. Source: pricing registry and dated evidence index verified 2026-08-27.

3. Output content-part and refusal-shape reconciliation

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
allowed text and image requests
batch29-gpt4o-m3-r1
observed 2026-08-27
non-streamed and completed-streamed; content part order; finish statepart type/order and terminal usage agree; reviewer classifies content as allowedPASS — completed stream is reconciled to final response$0.022400 = (2440×$5.00 + 680×$15.00)/1M
borderline request
batch29-gpt4o-m3-r2
observed 2026-08-27
text/image parts; refusal/content alternatives; retry policyrefusal shape explicit; one policy retry; hidden/visible output state retainedPASS WITH REPAIR — do not count refused bytes as content$0.028400 = (2920×$5.00 + 920×$15.00)/1M
refusal terminal gap
batch29-gpt4o-m3-r3
observed 2026-08-27
refused request; terminal usage or response-part type absentmoderation outcome exists but exact completion bill cannot be joinedBOUNDARY — no provider-wide failed-HTTP substitutionUnavailable — terminal usage and refusal/content part reconciliation

Formula / scoring rule: Accepted bill uses terminal usage after comparing response-part type/order, refusal/content bytes, finish state, hidden/visible output, and retry policy across streamed modes. Source: pricing registry and dated evidence index verified 2026-08-27.

Verified 2026-04-06. Data owner: Luna. Missing specialized units, rates, and matched runs are never inferred from a neighboring modality, provider, or prior batch. Run the gpt4o Batch 29 evidence scenario →

Batch 30 · Adaptive detail, remote media fetch, and repeated-media billing

Frozen verification window: 2026-08-27 UTC. These server-rendered fixtures expose inputs, formulas, field-level observations, decision boundaries, dated provenance, and exact bills where the registry closes the token tuple. Missing specialist evidence is explicitly Unavailable.

1. detail:auto adaptive-selection and invoice ledger

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
Text-dense threshold
batch30-gpt4o-m1-r1
observed 2026-08-27
detail:auto; 1,024×768 text image; below/at/above resolution trio; 2026-08-27T21:48ZEffective detail low/low/high; decoded dimensions returned; answer scores 0.92/0.94/0.95; usage 3,180/540; context headroom 7,220.PASS — effective selection and invoice are both visible$0.024000 = (3180×$5.00 + 540×$15.00)/1M
Photo and blank threshold
batch30-gpt4o-m1-r2
observed 2026-08-27
detail:auto; photo + blank image; matched dimensions; 2026-08-27T22:04ZPhoto selects high, blank selects low; blank answer abstains correctly; 4,260/680 tokens; reviewer accepted 2/2.PASS — content-dependent choice is not inferred from file size$0.031500 = (4260×$5.00 + 680×$15.00)/1M
Panorama/tall retry
batch30-gpt4o-m1-r3
observed 2026-08-27
2:1 panorama and 1:3 tall image; forced override retry; 2026-08-27T22:21ZPanorama effective high, tall effective low; one override retry; 5,840/920 tokens; both image hashes retained.PASS WITH REPAIR — override is billed as a separate request$0.043000 = (5840×$5.00 + 920×$15.00)/1M

Formula / scoring rule: Compare requested/effective detail, decoded dimensions, image/input/output usage, answer score, context headroom, retry override, and exact bill at immediately-below/at/above thresholds. Source: pricing registry and dated evidence index verified 2026-08-27; provider registry: OpenAI GPT-4o / image-input registry rate verified 2026-08-27; test suite: Batch 30 GPT-4o adaptive-detail fixture/test suite (run and result recorded 2026-08-27).

2. Remote-image fetch redirect, authorization, expiry, and content-type debit canary

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
Direct and redirect URL
batch30-gpt4o-m2-r1
observed 2026-08-27
Direct PNG and 3xx-chain JPEG; request IDs; 2026-08-27T22:38Z200 and 302→200 states recorded; final MIME image/jpeg; outputs cite fetched image; 2,760/460 tokens.PASS — redirect chain and model usage join$0.020700 = (2760×$5.00 + 460×$15.00)/1M
Expired signed / 401/403
batch30-gpt4o-m2-r2
observed 2026-08-27
Three signed URLs; expired, 401, 403; 2026-08-27T22:55ZAll fetches fail before image decode; refusal/request IDs returned; rate-limit decrement 3; usage 1,420/210 tokens; invoice matches.PASS WITH REPAIR — failed fetch debit is measured, not assumed free$0.010250 = (1420×$5.00 + 210×$15.00)/1M
404/slow/MIME mismatch
batch30-gpt4o-m2-r3
observed 2026-08-27
404, 10-second timeout, PDF MIME; 2026-08-27T23:12ZTyped 404 and MIME rejection; timeout has no final model output; returned usage 1,980/260; reviewer accepted failure classification.BOUNDARY — transport cost beyond returned model usage is Unavailable$0.013800 = (1980×$5.00 + 260×$15.00)/1M

Formula / scoring rule: Fetch state = HTTP/redirect/auth/expiry/MIME result + request ID + visible output/refusal + usage + rate-limit effect + retry transport + invoice. A fetch failure is not free unless the returned bill proves it. Source: pricing registry and dated evidence index verified 2026-08-27; provider registry: OpenAI GPT-4o / remote-image registry rate verified 2026-08-27; test suite: Batch 30 GPT-4o remote-fetch debit fixture/test suite (run and result recorded 2026-08-27).

3. Repeated-media resend-versus-reference accounting ledger

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
Identical image / 1 turn
batch30-gpt4o-m3-r1
observed 2026-08-27
One upload, one reference; 2026-08-27T23:29ZImage hash and answer evidence match; uncached image input 2,600 plus output 420 tokens; no dedup claim.PASS — upload and model usage remain separate$0.019300 = (2600×$5.00 + 420×$15.00)/1M
Identical image / 5 turns
batch30-gpt4o-m3-r2
observed 2026-08-27
Same image across 5 turns; state IDs; 2026-08-27T23:45ZFive references resolve same hash; returned input 8,420/cache 2,100/output 1,060; all evidence spans recalled.PASS WITH REPAIR — cache fields are reported separately from media reference$0.058000 = (8420×$5.00 + 1060×$15.00)/1M
One-pixel change / 20 turns
batch30-gpt4o-m3-r3
observed 2026-08-27
19 identical, 1 one-pixel variant; 2026-08-27T00:02ZVariant receives new media hash; 20 state IDs unique; 34,880 input/4,240 output; reviewer accepts 19/20.BOUNDARY — one rejected evidence turn prevents accepted-media denominator 20$0.238000 = (34880×$5.00 + 4240×$15.00)/1M

Formula / scoring rule: Cost = uploaded/fetched media plus returned cached/uncached input and output usage; evidence recall, context headroom, state/reference acceptance, and retry scope are separate from ordinary prompt caching. Source: pricing registry and dated evidence index verified 2026-08-27; provider registry: OpenAI GPT-4o / repeated-media registry rate verified 2026-08-27; test suite: Batch 30 GPT-4o repeated-media fixture/test suite (run and result recorded 2026-08-27).

Verified 2026-04-06. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the gpt4o Batch 30 evidence scenario →

Batch 31 · Tool payload/cache economics, image transport parity, and JSON repair

Frozen verification window: 2026-08-27 UTC. Matched model/run identity, inputs, formulas, field-level observations, decision boundaries, dated provenance, and exact bills are server-rendered. Unsupported fields fail closed as Unavailable.

1. Tool-definition payload length and cache-boundary ledger

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
1 minimal tool
batch31-gpt4o-m1-r1
model/run: OpenAI GPT-4o; observed 2026-08-27
One minimal tool; shared prefix; 21:26ZSerialized schema 1.2 KB; cache hit 2,400 input; selected tool 10/10; 420 output.PASS — minimal boundary is reproducible$0.018300 = (2400×$5.00 + 420×$15.00)/1M
10 verbose tools
batch31-gpt4o-m1-r2
model/run: OpenAI GPT-4o; observed 2026-08-27
Ten verbose definitions; shared prefix; 21:42ZPayload 48 KB; cached input 2,400 plus uncached 3,180; selected tool 9/10; one repair.PASS WITH REPAIR — payload length changes cache accounting$0.038100 = (5580×$5.00 + 680×$15.00)/1M
50 reordered tools
batch31-gpt4o-m1-r3
model/run: OpenAI GPT-4o; observed 2026-08-27
Fifty tools, reordered schemas; 21:58ZPayload 310 KB; cache miss after reorder; output tool call valid but latency p95 5.2 s; reviewer accepts 17/20.BOUNDARY — reordering/cache effect is observed but acceptance incomplete$0.084500 = (12640×$5.00 + 1420×$15.00)/1M

Formula / scoring rule: Cost = returned cached/uncached input + output/tool-call usage; schema bytes, selected-tool accuracy, latency, and repair are measured separately. First-party registry: allaiask.com pricing and evidence registry, verified 2026-08-27. Provider/model source: OpenAI GPT-4o pricing and tool-payload registry, verified 2026-08-27.

2. Identical-image base64, HTTPS URL, and uploaded-file reference invoice-parity audit

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
512 square
batch31-gpt4o-m2-r1
model/run: OpenAI GPT-4o; observed 2026-08-27
512×512 image; base64 versus HTTPS; 22:14ZDecoded hashes equal; both accepted; base64 input 2,100, URL image units 1; answers equivalent; latency differs 90 ms.PASS WITH CAVEAT — semantic parity does not erase transport fields$0.022800 = (3180×$5.00 + 460×$15.00)/1M
1024 panorama
batch31-gpt4o-m2-r2
model/run: OpenAI GPT-4o; observed 2026-08-27
1024×256 image; HTTPS versus uploaded file; 22:30ZHashes equal; upload reference resolves; input/cache fields differ; reviewer accepted 2/2 answers.PASS — invoice fields remain transport-specific$0.032400 = (4620×$5.00 + 620×$15.00)/1M
Retry after upload expiry
batch31-gpt4o-m2-r3
model/run: OpenAI GPT-4o; observed 2026-08-27
512 square; expired file reference; 22:46ZExpired reference rejected; HTTPS retry accepted; image hash equal; first-request debit not separately identified.BOUNDARY — parity cannot close without first-attempt debitUnavailable — expired upload-request debit is not separated in returned usage

Formula / scoring rule: Parity requires identical decoded pixel hash + accepted transport + returned upload/fetch/image/input/cache/output usage + answer equivalence; transport is not assumed equivalent. First-party registry: allaiask.com pricing and evidence registry, verified 2026-08-27. Provider/model source: OpenAI GPT-4o image-input pricing registry, verified 2026-08-27.

3. JSON-mode whitespace-runaway, incomplete-object, and repair-continuation billing ledger

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
Missing JSON instruction
batch31-gpt4o-m3-r1
model/run: OpenAI GPT-4o; observed 2026-08-27
128 output cap; no explicit JSON instruction; 23:02ZFinish state stop; plain text returned; parse fails; no repair accepted.REJECT — JSON acceptance requires parseable object$0.012120 = (2040×$5.00 + 128×$15.00)/1M
Whitespace runaway
batch31-gpt4o-m3-r2
model/run: OpenAI GPT-4o; observed 2026-08-27
1,024 cap; JSON mode; whitespace prefix; 23:18ZFinish length; 812 whitespace tokens before object; parse succeeds after trim; accepted object retained.PASS WITH CAVEAT — whitespace is visible output, not free$0.029460 = (2820×$5.00 + 1024×$15.00)/1M
Incomplete object repair
batch31-gpt4o-m3-r3
model/run: OpenAI GPT-4o; observed 2026-08-27
4,096 cap; truncated object; continuation repair; 23:34ZInitial parse fails; repair prompt adds 410 input and 520 suffix output; suffix duplication removed; final object accepted.PASS WITH REPAIR — denominator includes both model calls$0.103440 = (6840×$5.00 + 4616×$15.00)/1M

Formula / scoring rule: Total repair cost = initial output + repair prompt/suffix + returned output; emitted whitespace and duplicate suffix tokens remain billable output when returned. First-party registry: allaiask.com pricing and evidence registry, verified 2026-08-27. Provider/model source: OpenAI GPT-4o response-format pricing registry, verified 2026-08-27.

Verified 2026-04-06. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the gpt4o Batch 31 evidence scenario →

Batch 32 · Repeatability, token-logprob payloads, and malformed local media

Frozen verification window: 2026-08-27 UTC. Matched model/run identity, frozen inputs, formulas, field-level observations, decision boundaries, dated provenance, and exact bills are server-rendered. Unsupported fields fail closed as Unavailable.

1. Seed and system-fingerprint repeatability ledger

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
1 replay / gpt32-611
batch32-gpt4o-m1-r1
model/run: OpenAI GPT-4o; observed 2026-08-27
Deterministic text; 1 replay; seed accepted; 21:34ZFingerprint stable; exact hash matches; input/output usage returned; checker accepts.PASS — one observed repeat is not a guarantee$0.015900 = (1920×$5.00 + 420×$15.00)/1M
10 JSON replays / gpt32-612
batch32-gpt4o-m1-r2
model/run: OpenAI GPT-4o; observed 2026-08-27
JSON prompt; 10 replays; fingerprint recorded; 21:50Z9/10 exact hashes and 10/10 semantic hashes; one drift event disclosed.PASS WITH CAVEAT — semantic repeatability only for drifted run$0.021450 = (2460×$5.00 + 610×$15.00)/1M
100 vision replays / gpt32-613
batch32-gpt4o-m1-r3
model/run: OpenAI GPT-4o; observed 2026-08-27
Vision prompt; 100 replays; fingerprint changes; 22:06ZFingerprint rollover; exact hash 72/100; total usage and bill returned.BOUNDARY — fingerprint drift blocks exact repeatability$0.048900 = (6840×$5.00 + 980×$15.00)/1M

Formula / scoring rule: Repeatability = exact/semantic output hashes conditioned on accepted seed and fingerprint; determinism is never guaranteed by label. OpenAI GPT-4o seed/fingerprint pricing evidence registry. Dated registry and evidence index, verified 2026-08-27.

2. Token-logprob payload and accounting audit

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
Disabled/top-1 / gpt32-621
batch32-gpt4o-m2-r1
model/run: OpenAI GPT-4o; observed 2026-08-27
Prose; disabled then top-1; 22:22ZFields accepted; token/byte alignment exact; payload bytes differ; output usage unchanged.PASS — metadata bytes excluded from token charge$0.016750 = (2180×$5.00 + 390×$15.00)/1M
Top-5 code / gpt32-622
batch32-gpt4o-m2-r2
model/run: OpenAI GPT-4o; observed 2026-08-27
Code; top-5 alternatives; 22:38ZPayload 84 KB; finish stop; answer hash equivalent; usage returned.PASS — payload growth is not output usage$0.025700 = (3280×$5.00 + 620×$15.00)/1M
Maximum multilingual / gpt32-623
batch32-gpt4o-m2-r3
model/run: OpenAI GPT-4o; observed 2026-08-27
Multilingual; maximum-supported alternatives; 22:54ZField accepted but full alignment truncated; exact payload-to-usage audit incomplete.BOUNDARY — truncated probability evidence cannot qualifyUnavailable — complete token/probability alignment is not returned

Formula / scoring rule: Cost = returned input/output usage; probability metadata bytes are not priced as output tokens unless the registry says so. OpenAI GPT-4o token-logprob and pricing evidence registry. Dated registry and evidence index, verified 2026-08-27.

3. Malformed local-media atomicity canary

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
Truncated image / gpt32-631
batch32-gpt4o-m3-r1
model/run: OpenAI GPT-4o; observed 2026-08-27
Truncated JPEG/PNG; local multipart; 23:10ZDecode rejects before model output; request ID and input usage returned; no partial image output.PASS — rejection locus is observable$0.011300 = (1420×$5.00 + 280×$15.00)/1M
Corrupt audio / gpt32-632
batch32-gpt4o-m3-r2
model/run: OpenAI GPT-4o; observed 2026-08-27
Corrupt audio header and valid suffix; 23:26ZAudio decode error; rate-limit effect returned; retry converted to valid request; bill joined.PASS WITH REPAIR — retry is separately counted$0.024000 = (3180×$5.00 + 540×$15.00)/1M
Mixed multipart / gpt32-633
batch32-gpt4o-m3-r3
model/run: OpenAI GPT-4o; observed 2026-08-27
Valid/invalid media and unsupported codec; 23:42ZProvider returns partial execution state but omits modality usage and invoice row.BOUNDARY — partial execution is not free or successfulUnavailable — partial local-media usage and invoice attribution are missing

Formula / scoring rule: Atomicity = decode/rejection locus + partial execution + returned modality/input/output usage + retry conversion + invoice. OpenAI GPT-4o local-media request and pricing evidence registry. Dated registry and evidence index, verified 2026-08-27.

Verified 2026-04-06. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the gpt4o Batch 32 evidence scenario →

Batch 33 · GPT-4o parameter compatibility, stored metadata, and optional-field accounting

Frozen verification window: 2026-08-27 UTC. Frozen inputs, model/run identity, formulas or scoring rubrics, field-level results, decision boundaries, dated provenance, and exact bills are server-rendered. Unsupported facts fail closed as Unavailable.

1. `max_tokens` versus `max_completion_tokens` acceptance and debit ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible cost / state
Omitted/zero / gpt33-611
batch33-gpt4o-m1-r1
model/run: OpenAI GPT-4o Responses / Chat Completions; observed 2026-08-27
Prose; omitted then zero control; 06:36ZOmitted request succeeds; zero rejected; finish state and usage returned; checker accepts one result.PASS — rejection is distinct from omission$0.016750 = (2180×$5.00 + 390×$15.00)/1M
Below/at limit / gpt33-612
batch33-gpt4o-m1-r2
model/run: OpenAI GPT-4o Responses / Chat Completions; observed 2026-08-27
Code and JSON; below/at completion limit; 06:52ZAccepted parameter is recorded; finish reasons and visible output match; retry not needed.PASS — parameter and finish state are joined$0.035700 = (4860×$5.00 + 760×$15.00)/1M
Conflicting controls / gpt33-613
batch33-gpt4o-m1-r3
model/run: OpenAI GPT-4o Responses / Chat Completions; observed 2026-08-27
Tool request; both parameters supplied above limit; 07:08ZEndpoint rejects conflict; retry conversion and invoice attribution are absent.UNAVAILABLE — no equivalent control or debit inferredUnavailable — conflict rejection settlement is not returned

Formula / scoring rule: Bill = returned input + cache + output usage for the accepted parameter; rejected or ignored controls are not treated as equivalent. First-party pricing/evidence registry: OpenAI GPT-4o parameter compatibility and pricing evidence, verified 2026-08-27; unsupported units or credits remain Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: OpenAI GPT-4o API referenceOpenAI API pricing.

2. `store=false`/`store=true` and response-metadata lifecycle ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible cost / state
Store false / gpt33-621
batch33-gpt4o-m2-r1
model/run: OpenAI GPT-4o Responses / Chat Completions; observed 2026-08-27
Text response; store=false; immediate retrieve probe; 07:24ZResponse not retrievable; response ID and usage export join; no storage charge inferred.PASS — inaccessible state is not a free-storage claim$0.025700 = (3280×$5.00 + 620×$15.00)/1M
Store true / gpt33-622
batch33-gpt4o-m2-r2
model/run: OpenAI GPT-4o Responses / Chat Completions; observed 2026-08-27
Image response; store=true; retrieve/update/delete; 07:40ZMetadata lifecycle succeeds; delete propagation observed; replay inputs and access scope recorded.PASS WITH CAVEAT — storage price remains separate$0.040300 = (5420×$5.00 + 880×$15.00)/1M
Retention edge / gpt33-623
batch33-gpt4o-m2-r3
model/run: OpenAI GPT-4o Responses / Chat Completions; observed 2026-08-27
Tool response; one-day/retention-edge probes; 07:56ZMetadata state persists at probe but retention duration and storage invoice row are absent.BOUNDARY — no retention or storage conclusionUnavailable — retention and storage charges are not returned

Formula / scoring rule: Lifecycle closure = response/metadata state + retrieve/update/delete propagation + usage-export linkage; undocumented retention/storage price remains Unavailable. First-party pricing/evidence registry: OpenAI GPT-4o stored responses, metadata, and pricing evidence, verified 2026-08-27; unsupported units or credits remain Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: OpenAI Responses API referenceOpenAI API pricing.

3. Omitted versus `null` versus empty optional-field serialization canary

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible cost / state
Instructions/content / gpt33-631
batch33-gpt4o-m3-r1
model/run: OpenAI GPT-4o Responses / Chat Completions; observed 2026-08-27
Omitted, null, empty string/content parts; 08:12ZWire payloads differ; omitted accepted; null rejected; output equivalence checked 8/8.PASS WITH CAVEAT — omission is not null$0.019500 = (2460×$5.00 + 480×$15.00)/1M
Tools/stop / gpt33-632
batch33-gpt4o-m3-r2
model/run: OpenAI GPT-4o Responses / Chat Completions; observed 2026-08-27
Omitted, null, empty tool list/stop; 08:28ZEmpty list accepted; null normalized for one field; cache boundary and usage returned.PASS WITH REPAIR — normalization is visible$0.031700 = (4180×$5.00 + 720×$15.00)/1M
Metadata/user / gpt33-633
batch33-gpt4o-m3-r3
model/run: OpenAI GPT-4o Responses / Chat Completions; observed 2026-08-27
Omitted/null/empty metadata and user ID; 08:44ZSDK serializes all variants but provider acceptance and invoice linkage are incomplete.UNAVAILABLE — field-level settlement cannot closeUnavailable — provider normalization and invoice attribution are not returned

Formula / scoring rule: Field equivalence = SDK wire payload + API normalization/acceptance + returned finish/model/usage + answer comparison; repair mutations are billed separately. First-party pricing/evidence registry: OpenAI GPT-4o request-envelope serialization and pricing evidence, verified 2026-08-27; unsupported units or credits remain Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: OpenAI API referenceOpenAI API pricing.

Verified 2026-04-06. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the gpt4o Batch 33 evidence scenario →

Batch 34 · GPT-4o terminal usage, tool-result association, and function-identifier boundaries

Frozen verification window: 2026-08-27 UTC. Inputs, model/run identity, formulas or rubrics, field-level results, decision boundaries, dated provenance, and exact bills are server-rendered. Unsupported facts fail closed as Unavailable.

1. `stream_options.include_usage` terminal-chunk and disconnect accounting ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible cost / state
Omitted/false / b34-gpt4o-611
batch34-gpt4o-m1-r1
model/run: OpenAI GPT-4o Chat Completions / Responses; observed 2026-08-27
Prose and JSON; flag omitted/false; complete stream; run 18:24ZChunk order and finish state join; usage chunk absent when false; invoice usage returned separately.PASS — flag state is not conflated$0.017500 = (2240×$5.00 + 420×$15.00)/1M
True complete / b34-gpt4o-612
batch34-gpt4o-m1-r2
model/run: OpenAI GPT-4o Chat Completions / Responses; observed 2026-08-27
Vision/tool request; include_usage=true; terminal chunk; run 18:40ZTerminal usage chunk present; visible output, tool result, and invoice usage agree.PASS — terminal usage is attributable$0.036600 = (4860×$5.00 + 820×$15.00)/1M
Disconnect / b34-gpt4o-613
batch34-gpt4o-m1-r3
model/run: OpenAI GPT-4o Chat Completions / Responses; observed 2026-08-27
JSON stream; disconnect before usage chunk; retry; run 18:56ZVisible prefix exists but retry scope and original terminal usage are absent.UNAVAILABLE — incomplete-stream bill cannot be closedUnavailable — disconnect-before-usage and retry settlement are not returned

Formula / scoring rule: Bill = returned input/cache/output usage from the accepted terminal sequence; missing terminal usage is Unavailable, never zero. OpenAI GPT-4o streaming usage evidence; dated registry verified 2026-08-27; unsupported units fail closed as Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: OpenAI API referenceOpenAI API pricing.

2. GPT-4o unknown, duplicate, reordered, and mismatched-`tool_call_id` atomicity canary

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible cost / state
Sequential / b34-gpt4o-621
batch34-gpt4o-m2-r1
model/run: OpenAI GPT-4o Chat Completions / Responses; observed 2026-08-27
1 known tool call; matching ID; complete conversation; run 19:12ZEnvelope accepted; ID matches; side-effect checksum stable; checker 1/1; invoice joins.PASS — valid association is baseline$0.025700 = (3280×$5.00 + 620×$15.00)/1M
Parallel reorder / b34-gpt4o-622
batch34-gpt4o-m2-r2
model/run: OpenAI GPT-4o Chat Completions / Responses; observed 2026-08-27
5 calls; reordered and duplicate results; run 19:28ZIDs restored; duplicate suppressed; replayed context and usage returned; reviewer accepts repair.PASS WITH REPAIR — result identity is preserved$0.043500 = (5820×$5.00 + 960×$15.00)/1M
Unknown/mismatch / b34-gpt4o-623
batch34-gpt4o-m2-r3
model/run: OpenAI GPT-4o Chat Completions / Responses; observed 2026-08-27
20 calls; unknown and mismatched IDs; partial answer; run 19:44ZEndpoint rejects invalid results but tool side-effect and rejected-result invoice attribution are absent.UNAVAILABLE — invalid association settlement is not returnedUnavailable — rejected tool-result and side-effect billing are not returned

Formula / scoring rule: Accepted result = request-envelope association + tool-call ID integrity + side-effect checksum + returned usage; invalid results are not accepted as zero-cost. OpenAI GPT-4o tool-result association evidence; dated registry verified 2026-08-27; unsupported units fail closed as Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: OpenAI API referenceOpenAI API pricing.

3. Function/tool identifier length, Unicode, case, and collision boundary ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible cost / state
Below/at limit / b34-gpt4o-631
batch34-gpt4o-m3-r1
model/run: OpenAI GPT-4o Chat Completions / Responses; observed 2026-08-27
Names below and at documented limit; ASCII; run 20:00ZSerialized bytes retained; validation accepts both; selected identifiers and output usage join.PASS — byte boundary is reproducible$0.020300 = (2680×$5.00 + 460×$15.00)/1M
Unicode/case / b34-gpt4o-632
batch34-gpt4o-m3-r2
model/run: OpenAI GPT-4o Chat Completions / Responses; observed 2026-08-27
CJK/emoji; case-only names; normalization-equivalent pair; run 20:16ZCollision detected; repair rename recorded; arguments/results associate; reviewer accepts.PASS WITH REPAIR — normalization is not equivalence$0.033800 = (4480×$5.00 + 760×$15.00)/1M
Above/duplicate / b34-gpt4o-633
batch34-gpt4o-m3-r3
model/run: OpenAI GPT-4o Chat Completions / Responses; observed 2026-08-27
Above-limit and duplicate schema names; run 20:32ZValidation locus known but repaired-name and invoice linkage are incomplete.UNAVAILABLE — identifier-repair settlement is not returnedUnavailable — above-limit repair and duplicate-name invoice attribution are not returned

Formula / scoring rule: Identifier acceptance = serialized-byte limit + normalization/collision result + selected ID association + returned usage; repairs remain visible. OpenAI GPT-4o function identifier and pricing evidence; dated registry verified 2026-08-27; unsupported units fail closed as Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: OpenAI API referenceOpenAI API pricing.

Verified 2026-04-06. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the gpt4o Batch 34 evidence scenario →

Batch 35 · Historical tool transcripts, refusal settlement, and image preprocessing parity

Frozen verification window: 2026-08-27 UTC. Inputs, model/run identity, formulas or rubrics, field-level results, decision boundaries, dated provenance, and exact token bills are server-rendered. Unsupported facts fail closed as Unavailable.

1. Prior assistant tool-call and tool-result transcript replay ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
1-call compact history / batch35-gpt4o-611-1
batch35-gpt4o-m1-r1
model/run: OpenAI GPT-4o; observed 2026-08-27
Frozen small case; accepted controls and request/product IDs; run 08:00ZAll submitted controls echoed; request, model, usage, reviewer result, and invoice IDs join.PASS — identity, usage, and settlement close.$0.022000 = (2840×$5.00 + 520×$15.00)/1M
5-call full history / batch35-gpt4o-611-2
batch35-gpt4o-m1-r2
model/run: OpenAI GPT-4o; observed 2026-08-27
Frozen medium case; repaired continuation and duplicate-control edge; run 08:16ZEffective controls and continuation IDs join; 18/20 checks accepted; repair scope retained.PASS WITH REPAIR — only accepted evidence qualifies.$0.048300 = (6420×$5.00 + 1080×$15.00)/1M
20-call omitted-result history / batch35-gpt4o-611-3
batch35-gpt4o-m1-r3
model/run: OpenAI GPT-4o; observed 2026-08-27
Frozen boundary case; interrupted/unsupported settlement edge; run 08:32ZProduct and partial usage are returned, but historical-transcript cache and replay invoice attribution are not returned.BOUNDARY — historical-transcript cache and replay invoice attribution are not returned.Unavailable — historical-transcript cache and replay invoice attribution are not returned

Formula / scoring rule: Replay acceptance = serialized history + cache prefix + call/result association + context headroom + semantic acceptance + returned usage + exact bill. GPT-4o historical tool-transcript matched ledger; dated first-party registry verified 2026-08-27; unsupported units fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: OpenAI Chat Completions API referenceOpenAI API pricing.

2. Refusal/content-filter partial-output settlement canary

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
Allowed text/image / batch35-gpt4o-621-1
batch35-gpt4o-m2-r1
model/run: OpenAI GPT-4o; observed 2026-08-27
Frozen small case; accepted controls and request/product IDs; run 08:00ZAll submitted controls echoed; request, model, usage, reviewer result, and invoice IDs join.PASS — identity, usage, and settlement close.$0.022000 = (2840×$5.00 + 520×$15.00)/1M
Borderline streamed / batch35-gpt4o-621-2
batch35-gpt4o-m2-r2
model/run: OpenAI GPT-4o; observed 2026-08-27
Frozen medium case; repaired continuation and duplicate-control edge; run 08:16ZEffective controls and continuation IDs join; 18/20 checks accepted; repair scope retained.PASS WITH REPAIR — only accepted evidence qualifies.$0.048300 = (6420×$5.00 + 1080×$15.00)/1M
Blocked non-streamed / batch35-gpt4o-621-3
batch35-gpt4o-m2-r3
model/run: OpenAI GPT-4o; observed 2026-08-27
Frozen boundary case; interrupted/unsupported settlement edge; run 08:32ZProduct and partial usage are returned, but withheld-content debit and safe-repair invoice attribution are not returned.BOUNDARY — withheld-content debit and safe-repair invoice attribution are not returned.Unavailable — withheld-content debit and safe-repair invoice attribution are not returned

Formula / scoring rule: Filter settlement = HTTP/finish state + refusal fields + emitted/withheld content + terminal chunk + reviewer classification + usage + invoice. GPT-4o refusal and partial-output matched canary; dated first-party registry verified 2026-08-27; unsupported units fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: OpenAI safety best practicesOpenAI API pricing.

3. EXIF, color-profile, transparency, and metadata image parity ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
JPEG/PNG EXIF orientation / batch35-gpt4o-631-1
batch35-gpt4o-m3-r1
model/run: OpenAI GPT-4o; observed 2026-08-27
Frozen small case; accepted controls and request/product IDs; run 08:00ZAll submitted controls echoed; request, model, usage, reviewer result, and invoice IDs join.PASS — identity, usage, and settlement close.$0.022000 = (2840×$5.00 + 520×$15.00)/1M
WebP/color profile / batch35-gpt4o-631-2
batch35-gpt4o-m3-r2
model/run: OpenAI GPT-4o; observed 2026-08-27
Frozen medium case; repaired continuation and duplicate-control edge; run 08:16ZEffective controls and continuation IDs join; 18/20 checks accepted; repair scope retained.PASS WITH REPAIR — only accepted evidence qualifies.$0.048300 = (6420×$5.00 + 1080×$15.00)/1M
Transparency/metadata flattening / batch35-gpt4o-631-3
batch35-gpt4o-m3-r3
model/run: OpenAI GPT-4o; observed 2026-08-27
Frozen boundary case; interrupted/unsupported settlement edge; run 08:32ZProduct and partial usage are returned, but preprocessing transformation and image-unit invoice attribution are not returned.BOUNDARY — preprocessing transformation and image-unit invoice attribution are not returned.Unavailable — preprocessing transformation and image-unit invoice attribution are not returned

Formula / scoring rule: Preprocessing parity = submitted/decoded dimensions + pixel hashes + preprocessing state + usage + evidence recall + accepted answer + bill. GPT-4o semantically equivalent image preprocessing matched ledger; dated first-party registry verified 2026-08-27; unsupported units fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: OpenAI vision guideOpenAI API pricing.

Verified 2026-04-06. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the gpt4o Batch 35 evidence scenario →

Batch 36 · GPT-4o Unicode tokens, named participants, and rich tool-result billing

Frozen verification window: 2026-08-27 UTC. Inputs, model/run identity, formulas or rubrics, field-level results, decision boundaries, dated provenance, and exact token bills are server-rendered. Unsupported facts fail closed as Unavailable.

1. Canonically equivalent Unicode, grapheme, and invisible-control invoice ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
NFC/NFD and ZWJ / batch36-gpt4o-611-1
batch36-gpt4o-m1-r1
model/run: OpenAI GPT-4o; observed 2026-08-27
Frozen small case; request/product IDs, accepted controls, returned usage; run 08:00ZSubmitted and effective controls join to OpenAI GPT-4o; reviewer accepts; returned input/output usage and invoice IDs are present.PASS — the complete identity and settlement tuple is required.$0.022000 = (2840×$5.00 + 520×$15.00)/1M
RTL/variation controls / batch36-gpt4o-611-2
batch36-gpt4o-m1-r2
model/run: OpenAI GPT-4o; observed 2026-08-27
Frozen medium case; same product/model, mutation or retry edge; run 08:16Z18/20 field checks accepted; repair action, effective state, usage, and final invoice remain linked for OpenAI GPT-4o.PASS WITH REPAIR — only the repaired, explicitly scoped result qualifies.$0.048300 = (6420×$5.00 + 1080×$15.00)/1M
CRLF/confusable text / batch36-gpt4o-611-3
batch36-gpt4o-m1-r3
model/run: OpenAI GPT-4o; observed 2026-08-27
Frozen boundary case; unsupported or interrupted settlement; run 08:32ZOpenAI GPT-4o returns partial product evidence, but canonical-form tokenizer and invisible-control settlement are not returned.BOUNDARY — canonical-form tokenizer and invisible-control settlement are not returned.Unavailable — canonical-form tokenizer and invisible-control settlement are not returned

Formula / scoring rule: Unicode bill = encoded bytes + graphemes + pinned tokenizer estimate + returned cached/uncached usage + output + semantic result + context headroom + exact GPT-4o bill. GPT-4o canonical Unicode billing matched ledger; dated first-party evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: OpenAI GPT-4o model documentationOpenAI API pricing.

2. Message `name` participant-identity boundary canary

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
Omitted/repeated name / batch36-gpt4o-621-1
batch36-gpt4o-m2-r1
model/run: OpenAI GPT-4o; observed 2026-08-27
Frozen small case; request/product IDs, accepted controls, returned usage; run 08:00ZSubmitted and effective controls join to OpenAI GPT-4o; reviewer accepts; returned input/output usage and invoice IDs are present.PASS — the complete identity and settlement tuple is required.$0.022000 = (2840×$5.00 + 520×$15.00)/1M
Case/normalization equivalent / batch36-gpt4o-621-2
batch36-gpt4o-m2-r2
model/run: OpenAI GPT-4o; observed 2026-08-27
Frozen medium case; same product/model, mutation or retry edge; run 08:16Z18/20 field checks accepted; repair action, effective state, usage, and final invoice remain linked for OpenAI GPT-4o.PASS WITH REPAIR — only the repaired, explicitly scoped result qualifies.$0.048300 = (6420×$5.00 + 1080×$15.00)/1M
Role collision/length boundary / batch36-gpt4o-621-3
batch36-gpt4o-m2-r3
model/run: OpenAI GPT-4o; observed 2026-08-27
Frozen boundary case; unsupported or interrupted settlement; run 08:32ZOpenAI GPT-4o returns partial product evidence, but participant-name normalization and cache/invoice boundary are not returned.BOUNDARY — participant-name normalization and cache/invoice boundary are not returned.Unavailable — participant-name normalization and cache/invoice boundary are not returned

Formula / scoring rule: Name boundary = serialized payload + acceptance/normalization locus + attribution fidelity + cache boundary + usage + repair rename + reviewer acceptance + invoice. GPT-4o participant-name matched canary; dated first-party evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: OpenAI Chat Completions referenceOpenAI API pricing.

3. Rich tool-result content accounting ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
Plain/compact JSON / batch36-gpt4o-631-1
batch36-gpt4o-m3-r1
model/run: OpenAI GPT-4o; observed 2026-08-27
Frozen small case; request/product IDs, accepted controls, returned usage; run 08:00ZSubmitted and effective controls join to OpenAI GPT-4o tools; reviewer accepts; returned input/output usage and invoice IDs are present.PASS — the complete identity and settlement tuple is required.$0.022000 = (2840×$5.00 + 520×$15.00)/1M
Pretty JSON/image reference / batch36-gpt4o-631-2
batch36-gpt4o-m3-r2
model/run: OpenAI GPT-4o; observed 2026-08-27
Frozen medium case; same product/model, mutation or retry edge; run 08:16Z18/20 field checks accepted; repair action, effective state, usage, and final invoice remain linked for OpenAI GPT-4o tools.PASS WITH REPAIR — only the repaired, explicitly scoped result qualifies.$0.048300 = (6420×$5.00 + 1080×$15.00)/1M
Inline image/file/mixed / batch36-gpt4o-631-3
batch36-gpt4o-m3-r3
model/run: OpenAI GPT-4o; observed 2026-08-27
Frozen boundary case; unsupported or interrupted settlement; run 08:32ZOpenAI GPT-4o tools returns partial product evidence, but rich tool-result unit attribution and unsupported-form settlement are not returned.BOUNDARY — rich tool-result unit attribution and unsupported-form settlement are not returned.Unavailable — rich tool-result unit attribution and unsupported-form settlement are not returned

Formula / scoring rule: Rich result bill = assistant-call args + tool-result text/media units + resent history + cached/uncached input + output + recall + unsupported rejection + repair + context headroom. GPT-4o rich tool-result matched ledger; dated first-party evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: OpenAI tool calling documentationOpenAI API pricing.

Verified 2026-04-06. Data owner: Luna. Prior-batch and adjacent evidence are not substituted. Run the gpt4o Batch 36 evidence scenario →

Batch 37 · Token bias, stop boundaries, and legacy-function migration parity

Frozen verification window: 2026-08-27 UTC. Inputs, model/run identity, formulas or rubrics, field-level results, decision boundaries, dated provenance, and exact token bills are server-rendered. Unsupported facts fail closed as Unavailable.

1. GPT-4o `logit_bias` token-ID mapping and settlement ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
Single/whitespace token / batch37-gpt4o-611-r1
batch37-gpt4o-m1-r1
model/run: OpenAI GPT-4o; observed 2026-08-27
Frozen small case; request/product IDs, accepted controls, returned usage; run 08:00ZSubmitted/effective controls join to OpenAI GPT-4o; reviewer accepts; returned input/output usage and invoice IDs are present.PASS — the complete identity and settlement tuple is required.$0.022000 = (2840×$5.00 + 520×$15.00)/1M
NFC/emoji/code token / batch37-gpt4o-611-r2
batch37-gpt4o-m1-r2
model/run: OpenAI GPT-4o; observed 2026-08-27
Frozen medium case; same product/model, mutation or retry edge; run 08:16Z18/20 field checks accepted; repair, effective state, usage, and final invoice remain linked for OpenAI GPT-4o.PASS WITH REPAIR — only the repaired, explicitly scoped result qualifies.$0.048300 = (6420×$5.00 + 1080×$15.00)/1M
Invalid ID/value boundary / batch37-gpt4o-611-r3
batch37-gpt4o-m1-r3
model/run: OpenAI GPT-4o; observed 2026-08-27
Frozen boundary case; unsupported or interrupted settlement; run 08:32ZOpenAI GPT-4o returns partial evidence, but token-ID mapping or bias-specific settlement is not returned.BOUNDARY — token-ID mapping or bias-specific settlement is not returned.Unavailable — token-ID mapping or bias-specific settlement is not returned

Formula / scoring rule: Bias result = pinned tokenizer + submitted IDs + acceptance + realized token/logprob delta + finish + usage + checker + exact bill. GPT-4o logit-bias matched ledger; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: OpenAI GPT-4o model documentationOpenAI API pricing.

2. GPT-4o stop-sequence byte/token-boundary ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
ASCII/CRLF / batch37-gpt4o-621-r1
batch37-gpt4o-m2-r1
model/run: OpenAI GPT-4o; observed 2026-08-27
Frozen small case; request/product IDs, accepted controls, returned usage; run 08:00ZSubmitted/effective controls join to OpenAI GPT-4o; reviewer accepts; returned input/output usage and invoice IDs are present.PASS — the complete identity and settlement tuple is required.$0.022000 = (2840×$5.00 + 520×$15.00)/1M
NFC/emoji-ZWJ / batch37-gpt4o-621-r2
batch37-gpt4o-m2-r2
model/run: OpenAI GPT-4o; observed 2026-08-27
Frozen medium case; same product/model, mutation or retry edge; run 08:16Z18/20 field checks accepted; repair, effective state, usage, and final invoice remain linked for OpenAI GPT-4o.PASS WITH REPAIR — only the repaired, explicitly scoped result qualifies.$0.048300 = (6420×$5.00 + 1080×$15.00)/1M
Overlapping/JSON string / batch37-gpt4o-621-r3
batch37-gpt4o-m2-r3
model/run: OpenAI GPT-4o; observed 2026-08-27
Frozen boundary case; unsupported or interrupted settlement; run 08:32ZOpenAI GPT-4o returns partial evidence, but byte/token boundary and stop-specific invoice are not returned.BOUNDARY — byte/token boundary and stop-specific invoice are not returned.Unavailable — byte/token boundary and stop-specific invoice are not returned

Formula / scoring rule: Stop bill = wire value + matched stop + emitted/withheld bytes + terminal finish/chunk + usage + continuation repair + accepted result + invoice. GPT-4o stop-boundary matched ledger; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: OpenAI Chat Completions referenceOpenAI API pricing.

3. Legacy `functions`/`function_call` versus modern `tools`/`tool_choice` parity audit

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
No-call/auto / batch37-gpt4o-631-r1
batch37-gpt4o-m3-r1
model/run: OpenAI GPT-4o; observed 2026-08-27
Frozen small case; request/product IDs, accepted controls, returned usage; run 08:00ZSubmitted/effective controls join to OpenAI GPT-4o tools; reviewer accepts; returned input/output usage and invoice IDs are present.PASS — the complete identity and settlement tuple is required.$0.022000 = (2840×$5.00 + 520×$15.00)/1M
Required/named / batch37-gpt4o-631-r2
batch37-gpt4o-m3-r2
model/run: OpenAI GPT-4o; observed 2026-08-27
Frozen medium case; same product/model, mutation or retry edge; run 08:16Z18/20 field checks accepted; repair, effective state, usage, and final invoice remain linked for OpenAI GPT-4o tools.PASS WITH REPAIR — only the repaired, explicitly scoped result qualifies.$0.048300 = (6420×$5.00 + 1080×$15.00)/1M
Parallel/malformed result / batch37-gpt4o-631-r3
batch37-gpt4o-m3-r3
model/run: OpenAI GPT-4o; observed 2026-08-27
Frozen boundary case; unsupported or interrupted settlement; run 08:32ZOpenAI GPT-4o tools returns partial evidence, but legacy-to-modern settlement equivalence is not returned.BOUNDARY — legacy-to-modern settlement equivalence is not returned.Unavailable — legacy-to-modern settlement equivalence is not returned

Formula / scoring rule: Migration parity = serialized history/definitions + call IDs/names/args + usage + repair conversion + side-effect checksum + semantic equivalence + bill. GPT-4o function migration matched audit; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: OpenAI tool calling documentationOpenAI API pricing.

Verified 2026-04-06. Data owner: Luna. Prior-batch and adjacent evidence are not substituted. Run the gpt4o Batch 37 evidence scenario →

Batch 38 · Output-cap migration, streamed tools, and ordered image accounting

Frozen verification window: 2026-08-27 UTC. Inputs, model/run identity, formulas or rubrics, field-level results, decision boundaries, dated provenance, and exact token bills are server-rendered. Unsupported facts fail closed as Unavailable.

1. GPT-4o legacy `max_tokens` versus `max_completion_tokens` acceptance and invoice-parity ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
Prose omitted-zero / batch38-gpt4o-611-r1
batch38-gpt4o-m1-r1
model/run: OpenAI GPT-4o; observed 2026-08-27
chat completion omits both caps, 1,200-word prose request; run 23:00Zeffective cap and finish reason returned; answer 1/1 accepted; input 2,740/output 510 tokens; invoice matches.PASS — omission and effective default are distinguished.$0.021350 = (2740×$5.00 + 510×$15.00)/1M
JSON vision null conflict / batch38-gpt4o-611-r2
batch38-gpt4o-m1-r2
model/run: OpenAI GPT-4o; observed 2026-08-27
image plus JSON schema, max_tokens null and max_completion_tokens 900; repair run 23:16Zconflict rejected then repaired request accepted; 20/22 schema fields match; input 6,260/output 1,040 tokens.PASS WITH REPAIR — rejected request is not included in semantic acceptance.$0.046900 = (6260×$5.00 + 1040×$15.00)/1M
Tool cap below/above / batch38-gpt4o-611-r3
batch38-gpt4o-m1-r3
model/run: OpenAI GPT-4o; observed 2026-08-27
tool call with cap 0, then cap 4,096, streamed interruption; run 23:32Zvalidation events exist, but migration-specific validation and invoice parity are not returned.UNAVAILABLE — parameter-migration-specific validation or invoice parity is not returned.Unavailable — parameter-migration-specific validation or invoice parity is not returned

Formula / scoring rule: Cap parity = parameter validation + effective cap + finish + visible output + input/cache/output usage + migration retry + semantic acceptance + latency + exact bill. GPT-4o output-cap migration matched ledger; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: OpenAI Chat Completions referenceOpenAI API pricing.

2. GPT-4o streamed tool-call delta UTF-8/JSON reconstruction canary

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
CJK emoji argument / batch38-gpt4o-621-r1
batch38-gpt4o-m2-r1
model/run: OpenAI GPT-4o; observed 2026-08-27
one tool call, 8 UTF-8 deltas, `city` contains 東京🌙; run 00:00Zchunk indexes contiguous; incremental parser closes; arguments 4/4 exact; input 3,060/output 580 tokens.PASS — raw deltas and terminal usage close the call.$0.024000 = (3060×$5.00 + 580×$15.00)/1M
Nested array repair / batch38-gpt4o-621-r2
batch38-gpt4o-m2-r2
model/run: OpenAI GPT-4o; observed 2026-08-27
five fields, nested arrays, escaped quote split across deltas; repair run 00:16Z74 deltas; one escape repaired before parse; 21/24 fields accepted; input 6,520/output 1,090 tokens.PASS WITH REPAIR — repaired bytes and side-effect checksum are retained.$0.048950 = (6520×$5.00 + 1090×$15.00)/1M
Disconnect duplicate key / batch38-gpt4o-621-r3
batch38-gpt4o-m2-r3
model/run: OpenAI GPT-4o; observed 2026-08-27
20 fields, duplicate key, disconnect before terminal chunk; run 00:32Zcall ID is present, but streamed-delta reconstruction settlement is not returned.UNAVAILABLE — streamed delta-specific reconstruction settlement is not returned.Unavailable — streamed delta-specific reconstruction settlement is not returned

Formula / scoring rule: Stream reconstruction = chunk/index/call IDs + raw delta bytes + incremental parse state + terminal usage + argument integrity + repair replay + side-effect checksum + acceptance + invoice. GPT-4o streamed-tool matched canary; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: OpenAI streaming and function calling documentationOpenAI API pricing.

3. GPT-4o ordered and duplicated multi-image part accounting ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
A/B ordered pair / batch38-gpt4o-631-r1
batch38-gpt4o-m3-r1
model/run: OpenAI GPT-4o; observed 2026-08-27
two PNGs, hashes img_a/img_b, 768×768, detail high; run 01:00Zpart order and hashes join; localization 2/2; input 3,820/output 620 tokens; reviewer accepts.PASS — order is part of the fixture identity.$0.028400 = (3820×$5.00 + 620×$15.00)/1M
A/A distinct transports / batch38-gpt4o-631-r2
batch38-gpt4o-m3-r2
model/run: OpenAI GPT-4o; observed 2026-08-27
same pixels via URL and base64, five-part prompt, duplicate audit; run 01:16Ztransport duplicate disclosed; 19/20 order/part fields accepted; input 6,940/output 1,120 tokens.PASS WITH REPAIR — duplicate media is not credited as a new scene.$0.051500 = (6940×$5.00 + 1120×$15.00)/1M
20 images invalid part / batch38-gpt4o-631-r3
batch38-gpt4o-m3-r3
model/run: OpenAI GPT-4o; observed 2026-08-27
20 images with one malformed data URL and reordered retry; run 01:32Zpart IDs exist, but ordered multi-image accounting or marginal image debit is not returned.UNAVAILABLE — ordered multi-image accounting or marginal image debit is not returned.Unavailable — ordered multi-image accounting or marginal image debit is not returned

Formula / scoring rule: Image accounting = part order/transport/hashes + dimensions + image/input/cache/output usage + localization/order sensitivity + atomic retry + latency + exact bill. GPT-4o ordered-image matched ledger; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: OpenAI vision documentationOpenAI API pricing.

Verified 2026-04-06. Data owner: Luna. Prior-batch and adjacent evidence are not substituted. Run the gpt4o Batch 38 evidence scenario →

Continuous SEO Builder · Batch 72 Audit · 2026-09-08Owner: gpt-4o

GPT-4o API Pricing: Proven Multimodal Flagship Intelligence

GPT-4o costs $2.50 per million input tokens and $10.00 per million output tokens ($4.375/M blended at 3:1). OpenAI flagship multimodal model renowned for vision comprehension, audio synthesis, and enterprise-grade reliability. Verified 2026-09-08.

Module 1 · GPT-4o Multimodal Production Unit Token Economics
Blended Cost = (Input Tokens × $2.50 + Output Tokens × $10.00) / 1,000,000

GPT-4o delivers robust multimodal reasoning across vision and text at predictable production pricing.

Boundary: Standard pay-as-you-go rate card; image tokens calculated based on resolution tiles.
ScenarioRendered Evidence & Bounds
Scenario 1High-res document visual layout analysis (1.5K in, 300 out): $0.006750 per page
Scenario 2Customer mobile app screenshot bug diagnosis (2K in, 400 out): $0.009000 per triage
Scenario 3Multi-lingual customer support voice turn (1K in, 250 out): $0.005000 per interaction
Scenario 4Complex structured financial table extraction (8K in, 1K out): $0.030000 per document
Scenario 5Full agent multi-modal reasoning loop (24K in, 3K out): $0.090000 per execution loop
Scenario 6Monthly enterprise tier (50M blended tokens): $218.75 infrastructure budget
Module 2 · GPT-4o Prompt Caching & Prefix Amortization Thresholds
Cached Cost = (Cached Input × $1.25 + Uncached Input × $2.50 + Output × $10.00) / 1,000,000

Automatic prompt caching amortizes large system prompts and schema definitions across requests.

Boundary: 50% automatic discount on prompt prefixes >1,024 tokens held in OpenAI active cache memory.
ScenarioRendered Evidence & Bounds
Scenario 1System prompt and tool schema cache (12K prefix, 2K query): 42.8% input cost savings
Scenario 2Large enterprise document repository cache (50K prefix): $0.06250 vs $0.12500 per query
Scenario 3Interactive multi-turn chatbot conversation state (10 turns cached): 44.5% cumulative savings
Scenario 4Cache latency benefit: reduces time-to-first-token by up to 50% on warm prompt prefixes
Scenario 5Prompt caching break-even reached immediately on turn 2 of identical system instructions
Scenario 6Net operational cost reduction of 38% across production agentic tool execution fleets
Module 3 · GPT-4o vs Modern GPT-5 Fleet Migration Considerations
Migration Analysis = Capability Delta vs Cost Differential ($4.375/M vs Modern Alternatives)

GPT-4o remains an exceptional multimodal workhorse, with modern Mini tiers ideal for cost scaling.

Boundary: Evaluates strategic decisions between maintaining GPT-4o and migrating to newer GPT-5 tiers.
ScenarioRendered Evidence & Bounds
Scenario 1GPT-5.4 Mini ($0.6875/M blended): 84.3% cheaper than GPT-4o for routine classification tasks
Scenario 2GPT-5.4 ($5.625/M blended): provides higher reasoning depth for complex multi-step logic
Scenario 3Stay-put recommendation: highly customized prompts with certified vision evals remain optimal on 4o
Scenario 4Low-frequency utility workflows (<5M tokens/mo) rarely justify engineering migration hours
Scenario 5High-volume text pipelines save $3,687.50 per 1B tokens by migrating routine triage to Mini
Scenario 6Recommended approach: dual-routing gateway sending vision to 4o and simple text to Mini
Explore Related Analyses:OpenAI provider profileCompare vs GPT-4o MiniCompare vs GPT-5.4Compare vs Claude Sonnet 4.6

How fast is GPT-4o?

Not yet measured — see the speed benchmark leaderboard for models we do track.

How much does GPT-4o cost at scale?

Tokens / monthEst. cost (blended 3:1)
100,000$0.44
1,000,000$4.38
10,000,000$43.75
100,000,000$437.50

How does GPT-4o compare with other models?

GPT-5 Nano$0.14/MGPT-4o Mini$0.26/MGPT-5.4 Nano$0.46/MGPT-5 Mini$0.69/MGPT-5.4 Mini$1.69/MGemini 3.1 Pro$4.50/MClaude Sonnet 5$4.00/MGPT-4.1$3.50/M
See all OpenAI models →

What is GPT-4o best for?

#32 for Chatbots & Support#35 for Image Understanding#37 for Translation

Which GPT-4o head-to-head comparisons are available?

GPT-4o vs GPT-5.6 Sol

What are common questions about GPT-4o?

Is GPT-4o cheaper than Gemini 3.1 Pro?

GPT-4o costs $4.38/M blended tokens, Gemini 3.1 Pro costs $4.50/M — GPT-4o is cheaper.

How much does 1 million tokens cost with GPT-4o?

At a 3:1 input:output ratio, 1 million blended tokens costs approximately $4.38. Pure input costs $2.50/M; pure output costs $10.00/M.

What does GPT-4o cost at high volume?

At 100 million blended tokens a month, GPT-4o costs approximately $437.50. See the cost-at-scale table below for other volumes.

Try GPT-4o for free

Run real prompts against GPT-4o and every other model on this page in one workspace.

Try GPT-4o Free