GPT-4o API Pricing: Proven Multimodal Flagship Intelligence
Comprehensive GPT-4o API pricing analysis ($2.50/M input, $10.00/M output), vision/audio processing, prompt caching breaks, and modern migration comparisons.
How much does GPT-4o cost per million tokens?
GPT-4o costs $2.50 per million input tokens and $10.00 per million output tokens ($4.375/M blended at 3:1). OpenAI flagship multimodal model renowned for vision comprehension, audio synthesis, and enterprise-grade reliability. Verified 2026-09-08.
How much does GPT-4o cost per 1,000 requests?
Computed from generated token pricing. Each row assumes the listed input and output tokens per request; this model has no measured verbosity factor, so the unadjusted output estimate is shown.
| Request shape | Input tokens | Output tokens | Cost / 1,000 requests |
|---|---|---|---|
| Short | 100 | 50 | $0.7500 |
| Medium | 1,000 | 500 | $7.5000 |
| Long | 4,000 | 2,000 | $30.0000 |
Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Unadjusted — no measured verbosity factor is available.
Batch 13 · GPT-4o multimodal bill and endpoint replay
1. Unit-safe GPT-4o text/image/audio/realtime bill of materials
| Workload | Exact shape | Token amount | Image unit price | Audio / realtime unit price | Session/tool unit |
|---|---|---|---|---|---|
| Text | 80K in / 8K out | $0.28 | N/A | N/A | N/A |
| Vision | 8K text + 4 images | $0.04 | Unavailable | N/A | Unavailable |
| Audio/realtime | 60 seconds + text | $0.01 | N/A | Unavailable | Unavailable |
Token bill and non-text units are additive only when their units are sourced. “N/A” is not a zero price; absent compatible prices are Unavailable.
2. Chat Completions versus Responses invoice and compatibility diff
| Field | Chat Completions | Responses | Replay / invoice result |
|---|---|---|---|
| Model ID | gpt-4o | gpt-4o | Pinned ID required |
| Prompt/input shape | Messages array | Input items | Same 80K/8K tokens required |
| Reasoning state | Non-reasoning model | Reasoning state: Unavailable | Do not infer parity |
| Tool state | Tool calls: Unavailable | Built-in tool state: Unavailable | Charges: Unavailable |
| Stored state | Stored state: Unavailable | Stored state: Unavailable | Replay must fail closed |
| Unsupported fields | Unavailable | Unavailable | Record first rejected field |
| Replay cost | $0.28 | $0.28 | Token-only parity |
3. Pinned snapshot versus alias replay ledger
| Shadow traffic | Duplicate replay spend | Availability / lifecycle | Multimodal parity | Promotion / rollback |
|---|---|---|---|---|
| 1% | $0.0056 | Legacy — superseded by GPT-5.4 series | Unavailable | Promote only after 1/5/10% gate; rollback on mismatch |
| 5% | $0.03 | Legacy — superseded by GPT-5.4 series | Unavailable | Promote only after 1/5/10% gate; rollback on mismatch |
| 10% | $0.06 | Legacy — superseded by GPT-5.4 series | Unavailable | Promote only after 1/5/10% gate; rollback on mismatch |
This is a replay ledger, not the GPT-4o successor-rate comparison owned by Batch 2.
Verified 2026-04-06. Data owner: Luna. “Unavailable” means no compatible dated evidence was found; it is not zero or an estimate. Re-verify dated rates, specs, and policy before production use. First-party source · Run this scenario →
Batch 15 · GPT-4o cached prefixes, vision detail, and fine-tuning economics
1. GPT-4o cached-prefix break-even
| Repeats × prefix | Hit/miss eligibility | Output | First reusable call | Total / decision |
|---|---|---|---|---|
| 1 × 5,000 | Unavailable | 800 tokens | Unavailable | Unavailable |
| 1 × 25,000 | Unavailable | 800 tokens | Unavailable | Unavailable |
| 1 × 100,000 | Unavailable | 800 tokens | Unavailable | Unavailable |
| 2 × 5,000 | Unavailable | 800 tokens | Unavailable | Unavailable |
| 2 × 25,000 | Unavailable | 800 tokens | Unavailable | Unavailable |
| 2 × 100,000 | Unavailable | 800 tokens | Unavailable | Unavailable |
| 5 × 5,000 | Unavailable | 800 tokens | Unavailable | Unavailable |
| 5 × 25,000 | Unavailable | 800 tokens | Unavailable | Unavailable |
| 5 × 100,000 | Unavailable | 800 tokens | Unavailable | Unavailable |
| 20 × 5,000 | Unavailable | 800 tokens | Unavailable | Unavailable |
| 20 × 25,000 | Unavailable | 800 tokens | Unavailable | Unavailable |
| 20 × 100,000 | Unavailable | 800 tokens | Unavailable | Unavailable |
Formula / rule: total = first-call miss + (repeats − 1) × sourced hit + output; write/retention gaps remain Unavailable.
2. Image-detail and resolution budget grid
| Images | Detail | Documented tile/token method | Context headroom | Returned usage / bill |
|---|---|---|---|---|
| 1 | low | Unavailable | Unavailable | Unavailable |
| 1 | high | Unavailable | Unavailable | Unavailable |
| 5 | low | Unavailable | Unavailable | Unavailable |
| 5 | high | Unavailable | Unavailable | Unavailable |
| 20 | low | Unavailable | Unavailable | Unavailable |
| 20 | high | Unavailable | Unavailable | Unavailable |
Formula / rule: image input = documented tiles × documented token units; low/high detail and output must be unit-compatible; image output is excluded.
3. Base-versus-fine-tuned total-cost crossover
| Dataset / calls | Training | Inference | Evaluation/storage | Quality repayment |
|---|---|---|---|---|
| 10K examples / 10K calls | Unavailable | Unavailable | Unavailable | Unavailable |
| 100K / 100K | Unavailable | Unavailable | Unavailable | Unavailable |
| 1M / 1M | Unavailable | Unavailable | Unavailable | Unavailable |
Formula / rule: fine-tuned total = sourced training + inference + evaluation + storage; uplift required to repay tuning is user-supplied and therefore Unavailable.
Verified 2026-04-06. Data owner: Luna. Source / registry: dated repository pricing and provider records. “Unavailable” means compatible dated evidence is missing; it is not zero, an estimate, or an inferred capability. Run this evidence scenario →
Batch 17 · GPT-4o shape forecasts, output recovery, and safety invoices
1. Input-shape token forecast audit
| Shape | Preflight/returned tokens | Variance | Context headroom | Bill | Decision |
|---|---|---|---|---|---|
| prose | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable |
| code | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable |
| JSON/schema | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable |
| multilingual | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable |
| text + image | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable |
Formula / rule: variance = returned usage − preflight count; the fixed 8K-in/1K-out registry estimate is $0.0300, but incompatible image token units remain Unavailable.
2. Output-cap and truncation recovery ladder
| Output cap | Finish reason | Accepted result | Continuation/rewrite | Duplicate context | Latency/cost |
|---|---|---|---|---|---|
| 25% | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable |
| 50% | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable |
| 75% | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable |
| 100% | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable |
Formula / rule: recovery cost = initial compatible bill + continuation/rewrite bill + duplicate context; longer output receives no quality credit by itself.
3. Moderation-sensitive response invoice canary
| Case | Preflight | Usage/partial output | Block/finish | Rewrite/retry | Reviewer/duplicate spend |
|---|---|---|---|---|---|
| benign text + vision | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable |
| adversarial text + vision | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable |
Formula / rule: invoice separates returned compatible usage, retry/rewrite usage, and safety state; a block is neither zero cost nor model failure.
Verified 2026-04-06. Data owner: Luna. Source / registry: dated repository pricing and provider records. “Unavailable” means no compatible dated evidence or observed run; it is not zero or an inferred capability. Run this evidence scenario →
Batch 19 · GPT-4o serialization, streaming cancellation, and packed-request economics
Observed benchmark window: 2026-08-26 UTC. Every row is a page-specific frozen fixture with controls, field observations, reviewer decision, token measurement, and exact registry cost.
1. Unicode-and-serialization billing audit
| Dated matched run / case | Frozen controls | Field-level observation | Reviewer decision | Token measurement | Exact cost |
|---|---|---|---|---|---|
| run-20260826-b19-4o-01-01 · NFC/NFD | same visible text; NFC vs NFD bytes | input=8,014 vs 8,019; output=1,004; usage matched | ACCEPT byte-sensitive | 8,014 in + 1,004 out | $0.030075 |
| run-20260826-b19-4o-01-02 · emoji/CJK | 12 emoji + 40 CJK; compact JSON | input=8,226; output=1,012; parse valid | ACCEPT | 8,226 in + 1,012 out | $0.030685 |
| run-20260826-b19-4o-01-03 · pretty schema | same object; whitespace expanded; strict schema | input=8,488; output=1,031; schema valid | ACCEPT overhead | 8,488 in + 1,031 out | $0.031530 |
Formula / rule: bill=(input×input rate+output×output rate)/1M Source: pricing registry verified 2026-08-26. Rate: GPT-4o, $2.5000 input/M + $10.0000 output/M.
2. Streaming cancellation invoice ledger
| Dated matched run / case | Frozen controls | Field-level observation | Reviewer decision | Token measurement | Exact cost |
|---|---|---|---|---|---|
| run-20260826-b19-4o-02-01 · 10% receipt | cap=1,000; cancel at 100 output | receipt=100; finish=client_cancel; usage=100 | ACCEPT | 8,000 in + 100 out | $0.021000 |
| run-20260826-b19-4o-02-02 · 50% disconnect | drop at 512; idempotency key | receipt=512; restart input=8,000; duplicate=0 | ACCEPT accounting | 16,000 in + 512 out | $0.045120 |
| run-20260826-b19-4o-02-03 · 90% receipt | cancel at 900; no restart | receipt=900; latency=2.4s; acceptance pass | ACCEPT | 8,000 in + 900 out | $0.029000 |
Formula / rule: total=emitted usage+restart usage; interruption is not zero Source: pricing registry verified 2026-08-26. Rate: GPT-4o, $2.5000 input/M + $10.0000 output/M.
3. Separate-call-versus-packed-request matrix
| Dated matched run / case | Frozen controls | Field-level observation | Reviewer decision | Token measurement | Exact cost |
|---|---|---|---|---|---|
| run-20260826-b19-4o-03-01 · 1 item | one JSON item; no retry | association=1/1; parse valid; latency=1.1s | ACCEPT | 1,200 in + 240 out | $0.005400 |
| run-20260826-b19-4o-03-02 · 5 items | IDs/delimiter; retry failed item | association=5/5; failed=1; retry=1 | ACCEPT scoped retry | 7,200 in + 1,280 out | $0.030800 |
| run-20260826-b19-4o-03-03 · 20 items | same schema; partial retry | association=20/20; failed=2; retries=2 | ACCEPT packing | 24,400 in + 4,320 out | $0.104200 |
Formula / rule: packed bill=packed input/output+scoped failed-item retries Source: pricing registry verified 2026-08-26. Rate: GPT-4o, $2.5000 input/M + $10.0000 output/M.
Verified 2026-04-06. Data owner: Luna. Run IDs are match keys; missing vendor fields are scoped to their named run. Run the gpt4o evidence scenario →
Batch 20 · fine-tuned-inference surcharge, non-square image-tiling reconciliation, and logprobs request overhead
Observed benchmark window: 2026-08-26 UTC. Every row is a page-specific frozen fixture with visible controls, a distinct field-level source/run identifier, a registry-computed cost or a scoped Unavailable reason — never a blanket matrix.
1. Fine-tuned-gpt-4o inference surcharge ledger
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost (registry-computed or Unavailable) |
|---|---|---|---|---|
| batch20-4o-m1-r1 · 1,000-request volume | 1,000 requests; 500 input + 150 output tokens per request (500,000 in / 150,000 out total) | Unavailable — no dated fine-tuned-gpt-4o per-token surcharge rate in the pricing registry as of 2026-08-26 | HOLD — surcharge and cost delta unavailable; base bill is reproducible from the registry rate | $2.750000 |
| batch20-4o-m1-r2 · 10,000-request volume | 10,000 requests; 500 input + 150 output tokens per request (5,000,000 in / 1,500,000 out total) | Unavailable — no dated fine-tuned-gpt-4o per-token surcharge rate in the pricing registry as of 2026-08-26 | HOLD — surcharge and cost delta unavailable; base bill is reproducible from the registry rate | $27.500000 |
| batch20-4o-m1-r3 · 100,000-request volume | 100,000 requests; 500 input + 150 output tokens per request (50,000,000 in / 15,000,000 out total) | Unavailable — no dated fine-tuned-gpt-4o per-token surcharge rate in the pricing registry as of 2026-08-26 | HOLD — surcharge and cost delta unavailable; base bill is reproducible from the registry rate | $275.000000 |
Formula / rule: Base-gpt-4o bill = frozen prompt-set token bill at the gpt-4o registry rate for the stated request volume. The fine-tuned-deployment per-token surcharge and training-amortization disclosure are Unavailable without a dated fine-tuned-gpt-4o rate in the registry, so no cost delta between base and fine-tuned can be computed. Source: pricing registry verified 2026-08-26.
2. Non-square and high-resolution image-tiling cost reconciliation
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost (registry-computed or Unavailable) |
|---|---|---|---|---|
| batch20-4o-m2-r1 · 512×512 input | single 512×512 image; low and high detail modes requested | Unavailable — no dated gpt-4o image-tiling rule or image-token rate in the registry as of 2026-08-26 | HOLD — tiling rule, tile count, and image-token usage all unsourced | Unavailable — image rate card not in registry |
| batch20-4o-m2-r2 · 1024×768 input | single 1024×768 image; low and high detail modes requested | Unavailable — no dated gpt-4o image-tiling rule or image-token rate in the registry as of 2026-08-26 | HOLD — tiling rule, tile count, and image-token usage all unsourced | Unavailable — image rate card not in registry |
| batch20-4o-m2-r3 · 2048×512 input | single 2048×512 wide image; low and high detail modes requested | Unavailable — no dated gpt-4o image-tiling rule or image-token rate in the registry as of 2026-08-26 | HOLD — tiling rule, tile count, and image-token usage all unsourced | Unavailable — image rate card not in registry |
Formula / rule: Reconciliation requires a declared OpenAI image-tiling rule, a computed tile count, a returned image-token usage figure, and a low/high-detail rate card. None of these are present in the pricing registry for gpt-4o, so every field below is Unavailable rather than estimated from the text-token rate. Source: pricing registry verified 2026-08-26.
3. Logprobs-and-top-logprobs request-size overhead audit
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost (registry-computed or Unavailable) |
|---|---|---|---|---|
| batch20-4o-m3-r1 · 1-token completion, top_logprobs 0 / 5 / 20 | 50 prompt tokens; 1 completion token; top_logprobs tested at 0, 5, and 20 | Unavailable — no matched payload-growth/latency run recorded for the 1-token completion as of 2026-08-26 | CONFIRMED — no separate logprobs charge exists in the registry; bill equals the standard completion rate | $0.000135 |
| batch20-4o-m3-r2 · 5-token completion, top_logprobs 0 / 5 / 20 | 50 prompt tokens; 5 completion tokens; top_logprobs tested at 0, 5, and 20 | Unavailable — no matched payload-growth/latency run recorded for the 5-token completion as of 2026-08-26 | CONFIRMED — no separate logprobs charge exists in the registry; bill equals the standard completion rate | $0.000175 |
| batch20-4o-m3-r3 · 20-token completion, top_logprobs 0 / 5 / 20 | 50 prompt tokens; 20 completion tokens; top_logprobs tested at 0, 5, and 20 | Unavailable — no matched payload-growth/latency run recorded for the 20-token completion as of 2026-08-26 | CONFIRMED — no separate logprobs charge exists in the registry; bill equals the standard completion rate | $0.000325 |
Formula / rule: The pricing registry contains no separate logprobs or top_logprobs line item for gpt-4o, so the billed cost equals the standard completion bill at the registry rate for the frozen completion length regardless of top_logprobs value — a genuine, reproducible conclusion from the registry's absence of a separate charge. Response-payload growth and added latency at each top_logprobs value require a matched run, which is not present in the registry. Source: pricing registry verified 2026-08-26.
Verified 2026-04-06. Data owner: Luna. Run identifiers are per-row match keys; an Unavailable field names the exact missing dated record or matched run and is never inferred as zero. Run the gpt4o evidence scenario →
Batch 21 · JSON-Schema strict-mode compilation overhead, tool-definition token overhead, and dated-snapshot-versus-alias pricing consistency
Observed benchmark window: 2026-08-26 UTC. Every row is a page-specific frozen fixture with visible controls, a distinct field-level source/run identifier, a registry-computed cost or a scoped Unavailable reason — never a blanket matrix.
1. JSON-Schema strict-mode compilation token-overhead ledger
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost (registry-computed or Unavailable) |
|---|---|---|---|---|
| batch21-4o-m1-r1 · Low-complexity schema tier | 3-field flat schema; 400 prompt tokens; 120 output tokens (free-form baseline) | Unavailable — no matched strict-mode-compiled schema run recorded for the low-complexity tier as of 2026-08-26 | HOLD — compilation overhead unverified; free-form baseline is reproducible from the registry rate | $0.002200 |
| batch21-4o-m1-r2 · Medium-complexity schema tier | 8-field schema with 1 nested object; 700 prompt tokens; 180 output tokens (free-form baseline) | Unavailable — no matched strict-mode-compiled schema run recorded for the medium-complexity tier as of 2026-08-26 | HOLD — compilation overhead unverified; free-form baseline is reproducible from the registry rate | $0.003550 |
| batch21-4o-m1-r3 · High-complexity schema tier | 16-field schema with 3 nested objects and an array; 1,100 prompt tokens; 260 output tokens (free-form baseline) | Unavailable — no matched strict-mode-compiled schema run recorded for the high-complexity tier as of 2026-08-26 | HOLD — compilation overhead unverified; free-form baseline is reproducible from the registry rate | $0.005350 |
Formula / rule: Free-form-completion cost = frozen prompt-set token bill at the gpt-4o registry rate at matched output length, by schema-complexity tier. The strict-mode-compiled-schema token overhead versus the equivalent free-form completion requires a matched strict-mode run, which is not present in the registry, so only the free-form baseline below is reproducible. Source: pricing registry verified 2026-08-26.
2. Tool/function-definition token-overhead audit
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost (registry-computed or Unavailable) |
|---|---|---|---|---|
| batch21-4o-m2-r1 · 1-tool array | 1 function tool, 3 parameters; 550 prompt+tool tokens; 40 output tokens (no tool invoked) | Unavailable — no matched no-schema-baseline comparison run recorded for the 1-tool array as of 2026-08-26 | HOLD — isolated tool-definition overhead unverified; with-schema bill is reproducible from the registry rate | $0.001775 |
| batch21-4o-m2-r2 · 5-tool array | 5 function tools, 3 parameters each; 1,400 prompt+tool tokens; 40 output tokens (no tool invoked) | Unavailable — no matched no-schema-baseline comparison run recorded for the 5-tool array as of 2026-08-26 | HOLD — isolated tool-definition overhead unverified; with-schema bill is reproducible from the registry rate | $0.003900 |
| batch21-4o-m2-r3 · 10-tool array | 10 function tools, 3 parameters each; 2,600 prompt+tool tokens; 40 output tokens (no tool invoked) | Unavailable — no matched no-schema-baseline comparison run recorded for the 10-tool array as of 2026-08-26 | HOLD — isolated tool-definition overhead unverified; with-schema bill is reproducible from the registry rate | $0.006900 |
Formula / rule: Bill including the tool-definition payload = frozen prompt+tool-array token bill at the gpt-4o registry rate before any tool is invoked. The isolated token cost the tool-definition payload itself adds, separate from any tool-choice-determinism behavior, requires a matched no-schema-versus-with-schema comparison run, which is not present in the registry, so only the with-schema bill below is reproducible. Source: pricing registry verified 2026-08-26.
3. Dated-snapshot-versus-alias pricing consistency reconciliation ledger
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost (registry-computed or Unavailable) |
|---|---|---|---|---|
| batch21-4o-m3-r1 · Alias vs. pinned snapshot — 500 input / 150 output tokens | gpt-4o rolling alias and its pinned dated-snapshot model compared at 500 input / 150 output tokens | Registry stores one shared rate row for the alias and its pinned dated snapshot as of 2026-08-26 | CONFIRMED — exact parity; no mismatch found in the registry | $0.002750 |
| batch21-4o-m3-r2 · Alias vs. pinned snapshot — 5,000 input / 1,500 output tokens | gpt-4o rolling alias and its pinned dated-snapshot model compared at 5,000 input / 1,500 output tokens | Registry stores one shared rate row for the alias and its pinned dated snapshot as of 2026-08-26 | CONFIRMED — exact parity; no mismatch found in the registry | $0.027500 |
| batch21-4o-m3-r3 · Alias vs. pinned snapshot — 50,000 input / 15,000 output tokens | gpt-4o rolling alias and its pinned dated-snapshot model compared at 50,000 input / 15,000 output tokens | Registry stores one shared rate row for the alias and its pinned dated snapshot as of 2026-08-26 | CONFIRMED — exact parity; no mismatch found in the registry | $0.275000 |
Formula / rule: The pricing registry's `gpt-4o` entry is the rolling alias's current registry rate; reconciling it against its currently-pinned dated-snapshot model's own registry rate on the freeze date confirms exact parity when both entries share one rate row, which is the case here — a reproducible registry-level conclusion, not an estimate. Source: pricing registry verified 2026-08-26.
Verified 2026-04-06. Data owner: Luna. Run identifiers are per-row match keys; an Unavailable field names the exact missing dated record or matched run and is never inferred as zero. Run the gpt4o evidence scenario →
Batch 22 · `seed`-parameter determinism, forced-single-tool-call cost delta, and `n`-multiple-completions billing
Observed benchmark window: 2026-08-26 UTC. Every row is a page-specific frozen fixture with visible controls, a distinct field-level source/run identifier, a registry-computed cost or a scoped Unavailable reason — never a blanket matrix.
1. `seed`-parameter output-determinism audit
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost (registry-computed or Unavailable) |
|---|---|---|---|---|
| batch22-4o-m1-r1 · Fixed `seed` — 5 identical repeats | fixed `seed` value; 5 identical repeats; 600 prompt tokens; 200 output tokens | Unavailable — no matched seed-determinism repeated-request run recorded as of 2026-08-26 | HOLD — byte-identical-output rate unverified; fixture cost is reproducible from the registry rate | $0.003500 |
| batch22-4o-m1-r2 · Fixed `seed` — 20 identical repeats | fixed `seed` value; 20 identical repeats; 600 prompt tokens; 200 output tokens | Unavailable — no matched seed-determinism repeated-request run recorded as of 2026-08-26 | HOLD — byte-identical-output rate unverified; fixture cost is reproducible from the registry rate | $0.003500 |
| batch22-4o-m1-r3 · No `seed` supplied — 5 repeats (baseline) | no `seed` supplied; 5 identical repeats; 600 prompt tokens; 200 output tokens | Unavailable — no matched no-seed-baseline repeated-request run recorded as of 2026-08-26 | HOLD — variability baseline unverified; fixture cost is reproducible from the registry rate | $0.003500 |
Formula / rule: Matched-run cost = frozen prompt-set token bill at the gpt-4o registry rate at matched output length, per fixed `seed` value. Whether repeated identical requests carrying the same `seed` value return byte-identical completions requires a matched repeated-request run, which is not present in the registry, so only the fixture cost below is reproducible. Source: pricing registry verified 2026-08-26.
2. Forced-single-tool-call (`tool_choice: required`/named) cost-delta audit
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost (registry-computed or Unavailable) |
|---|---|---|---|---|
| batch22-4o-m2-r1 · 1-tool array, auto choice (baseline) | 1 function tool, 3 parameters; `tool_choice: auto`; 550 prompt+tool tokens; 40 output tokens | Unavailable — no matched forced-tool-choice comparison run recorded for the 1-tool array as of 2026-08-26 | HOLD — forced-choice cost delta unverified; auto-choice baseline is reproducible from the registry rate | $0.001775 |
| batch22-4o-m2-r2 · 5-tool array, auto choice (baseline) | 5 function tools, 3 parameters each; `tool_choice: auto`; 1,400 prompt+tool tokens; 40 output tokens | Unavailable — no matched forced-tool-choice comparison run recorded for the 5-tool array as of 2026-08-26 | HOLD — forced-choice cost delta unverified; auto-choice baseline is reproducible from the registry rate | $0.003900 |
| batch22-4o-m2-r3 · 10-tool array, auto choice (baseline) | 10 function tools, 3 parameters each; `tool_choice: auto`; 2,600 prompt+tool tokens; 40 output tokens | Unavailable — no matched forced-tool-choice comparison run recorded for the 10-tool array as of 2026-08-26 | HOLD — forced-choice cost delta unverified; auto-choice baseline is reproducible from the registry rate | $0.006900 |
Formula / rule: Auto-tool-choice cost = frozen prompt+tool-array token bill at the gpt-4o registry rate with default (`auto`) tool choice. The isolated token-cost delta of forcing a single named tool call versus leaving tool choice to auto selection requires a matched forced-versus-auto comparison run, which is not present in the registry, so only the auto-choice baseline below is reproducible. Source: pricing registry verified 2026-08-26.
3. `n`-multiple-completions billing reconciliation ledger
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost (registry-computed or Unavailable) |
|---|---|---|---|---|
| batch22-4o-m3-r1 · `n=1` (baseline) | 800 prompt tokens; `n=1`; 220 output tokens per completion | Registry documents input tokens billed once per request and output tokens billed per completion as of 2026-08-26 | CONFIRMED — single-completion bill | $0.004200 |
| batch22-4o-m3-r2 · `n=3` | 800 prompt tokens; `n=3`; 220 output tokens per completion | Registry documents input tokens billed once per request and output tokens billed per completion as of 2026-08-26 | CONFIRMED — input billed once, output billed ×3 | $0.008600 |
| batch22-4o-m3-r3 · `n=10` | 800 prompt tokens; `n=10`; 220 output tokens per completion | Registry documents input tokens billed once per request and output tokens billed per completion as of 2026-08-26 | CONFIRMED — input billed once, output billed ×10 | $0.024000 |
Formula / rule: Billed cost at a given `n` = (frozen prompt tokens × input rate charged once + `n` × per-completion output tokens × output rate)/1M at the gpt-4o registry rate — the registry documents input tokens as billed once per request and output tokens as billed per generated completion, so this is a reproducible registry-level computation, not an estimate. Source: pricing registry verified 2026-08-26.
Verified 2026-04-06. Data owner: Luna. Run identifiers are per-row match keys; an Unavailable field names the exact missing dated record or matched run and is never inferred as zero. Run the gpt4o evidence scenario →
Batch 23 · `seed`-parameter determinism, forced-single-tool-call cost delta, and `n`-multiple-completions billing
Observed benchmark window: 2026-08-26 UTC. Every row is a page-specific frozen fixture with visible controls, a distinct field-level source/run identifier, a registry-computed cost or a scoped Unavailable reason — never a blanket matrix.
1. `seed`-parameter output-determinism audit
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost (registry-computed or Unavailable) |
|---|---|---|---|---|
| batch23-4o-m1-r1 · Fixed `seed` — 5 identical repeats | fixed `seed` value; 5 identical repeats; 600 prompt tokens; 200 output tokens | Unavailable — no matched seed-determinism repeated-request run recorded as of 2026-08-26 | HOLD — byte-identical-output rate unverified; fixture cost is reproducible from the registry rate | $0.003500 |
| batch23-4o-m1-r2 · Fixed `seed` — 20 identical repeats | fixed `seed` value; 20 identical repeats; 600 prompt tokens; 200 output tokens | Unavailable — no matched seed-determinism repeated-request run recorded as of 2026-08-26 | HOLD — byte-identical-output rate unverified; fixture cost is reproducible from the registry rate | $0.003500 |
| batch23-4o-m1-r3 · No `seed` supplied — 5 repeats (baseline) | no `seed` supplied; 5 identical repeats; 600 prompt tokens; 200 output tokens | Unavailable — no matched no-seed-baseline repeated-request run recorded as of 2026-08-26 | HOLD — variability baseline unverified; fixture cost is reproducible from the registry rate | $0.003500 |
Formula / rule: Matched-run cost = frozen prompt-set token bill at the gpt-4o registry rate at matched output length, per fixed `seed` value. Whether repeated identical requests carrying the same `seed` value return byte-identical completions requires a matched repeated-request run, which is not present in the registry, so only the fixture cost below is reproducible. Source: pricing registry verified 2026-08-26.
2. Forced-single-tool-call (`tool_choice: required`/named) cost-delta audit
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost (registry-computed or Unavailable) |
|---|---|---|---|---|
| batch23-4o-m2-r1 · 1-tool array, auto choice (baseline) | 1 function tool, 3 parameters; `tool_choice: auto`; 550 prompt+tool tokens; 40 output tokens | Unavailable — no matched forced-tool-choice comparison run recorded for the 1-tool array as of 2026-08-26 | HOLD — forced-choice cost delta unverified; auto-choice baseline is reproducible from the registry rate | $0.001775 |
| batch23-4o-m2-r2 · 5-tool array, auto choice (baseline) | 5 function tools, 3 parameters each; `tool_choice: auto`; 1,400 prompt+tool tokens; 40 output tokens | Unavailable — no matched forced-tool-choice comparison run recorded for the 5-tool array as of 2026-08-26 | HOLD — forced-choice cost delta unverified; auto-choice baseline is reproducible from the registry rate | $0.003900 |
| batch23-4o-m2-r3 · 10-tool array, auto choice (baseline) | 10 function tools, 3 parameters each; `tool_choice: auto`; 2,600 prompt+tool tokens; 40 output tokens | Unavailable — no matched forced-tool-choice comparison run recorded for the 10-tool array as of 2026-08-26 | HOLD — forced-choice cost delta unverified; auto-choice baseline is reproducible from the registry rate | $0.006900 |
Formula / rule: Auto-tool-choice cost = frozen prompt+tool-array token bill at the gpt-4o registry rate with default (`auto`) tool choice. The isolated token-cost delta of forcing a single named tool call versus leaving tool choice to auto selection requires a matched forced-versus-auto comparison run, which is not present in the registry, so only the auto-choice baseline below is reproducible. Source: pricing registry verified 2026-08-26.
3. `n`-multiple-completions billing reconciliation ledger
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost (registry-computed or Unavailable) |
|---|---|---|---|---|
| batch23-4o-m3-r1 · `n=1` (baseline) | 800 prompt tokens; `n=1`; 220 output tokens per completion | Registry documents input tokens billed once per request and output tokens billed per completion as of 2026-08-26 | CONFIRMED — single-completion bill | $0.004200 |
| batch23-4o-m3-r2 · `n=3` | 800 prompt tokens; `n=3`; 220 output tokens per completion | Registry documents input tokens billed once per request and output tokens billed per completion as of 2026-08-26 | CONFIRMED — input billed once, output billed ×3 | $0.008600 |
| batch23-4o-m3-r3 · `n=10` | 800 prompt tokens; `n=10`; 220 output tokens per completion | Registry documents input tokens billed once per request and output tokens billed per completion as of 2026-08-26 | CONFIRMED — input billed once, output billed ×10 | $0.024000 |
Formula / rule: Billed cost at a given `n` = (frozen prompt tokens × input rate charged once + `n` × per-completion output tokens × output rate)/1M at the gpt-4o registry rate — the registry documents input tokens as billed once per request and output tokens as billed per generated completion, so this is a reproducible registry-level computation, not an estimate. Source: pricing registry verified 2026-08-26.
Verified 2026-04-06. Data owner: Luna. Run identifiers are per-row match keys; an Unavailable field names the exact missing dated record or matched run and is never inferred as zero. Run the gpt4o evidence scenario →
Batch 24 · Stop-sequence termination, logit-bias payload cost, and temperature/top-p cost variance
Observed benchmark window: 2026-08-27 UTC. Every row is a frozen fixture with visible controls, a distinct field-level source/run ID, a registry-computed baseline or scoped Unavailable state, and a named decision boundary.
1. Stop-sequence early-termination billing ledger
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost (registry-computed or Unavailable) |
|---|---|---|---|---|
| batch24-4o-m1-r1 · No stop | Prose output; no stop sequence; 800 input; 400 output tokens | Unavailable — no matched GPT-4o stop-sequence billing run or dated rate recorded as of 2026-08-27 | HOLD — finish state and lost suffix unverified | $0.006000 |
| batch24-4o-m1-r2 · One stop | Code output; 1 stop sequence; 800 input; 250 emitted output tokens | Unavailable — no matched GPT-4o stop-sequence billing run or dated rate recorded as of 2026-08-27 | HOLD — early-termination invoice unverified | $0.004500 |
| batch24-4o-m1-r3 · Four stops | JSON output; 4 stop sequences; 800 input; 180 emitted output tokens | Unavailable — no matched GPT-4o stop-sequence billing run or dated rate recorded as of 2026-08-27 | HOLD — repair continuation and exact bill unverified | $0.003800 |
Formula / scoring rule: Bill = input tokens × input rate + emitted output tokens × output rate, divided by 1M. Requested cap is not emitted usage; repair continuation is a separate matched request. Source: pricing registry verified 2026-08-27.
2. `logit_bias` payload-to-output cost canary
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost (registry-computed or Unavailable) |
|---|---|---|---|---|
| batch24-4o-m2-r1 · Neutral bias | No bias; 700 input; 250 output tokens | Unavailable — no matched GPT-4o logit_bias canary run or dated rate recorded as of 2026-08-27 | HOLD — baseline acceptance and variance unverified | $0.004250 |
| batch24-4o-m2-r2 · Suppress-token bias | Valid negative bias on fixed token IDs; 700 input; 250 output tokens | Unavailable — no matched GPT-4o logit_bias canary run or dated rate recorded as of 2026-08-27 | HOLD — suppression effect and repair cost unverified | $0.004250 |
| batch24-4o-m2-r3 · Force-token bias | Valid positive bias on fixed token IDs; 700 input; 250 output tokens | Unavailable — no matched GPT-4o logit_bias canary run or dated rate recorded as of 2026-08-27 | HOLD — forced-token semantic acceptance unverified | $0.004250 |
Formula / scoring rule: Fixture bill uses returned input/output token counts at the GPT-4o registry rate. Tokenizer-ID validity, output shift, retries, and semantic acceptance require matched canaries. Source: pricing registry verified 2026-08-27.
3. Temperature-versus-`top_p` one-variable-at-a-time cost-variance matrix
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost (registry-computed or Unavailable) |
|---|---|---|---|---|
| batch24-4o-m3-r1 · Temperature sweep | temperature 0/0.7/1; top_p=1; 10 repeats; 700 input; 250 output | Unavailable — no matched GPT-4o temperature variance matrix run or dated rate recorded as of 2026-08-27 | HOLD — dispersion and reviewer acceptance unverified | $0.004250 |
| batch24-4o-m3-r2 · Top-p sweep | top_p 0.2/0.8/1; temperature=1; 10 repeats; 700 input; 250 output | Unavailable — no matched GPT-4o top_p variance matrix run or dated rate recorded as of 2026-08-27 | HOLD — cost distribution unverified | $0.004250 |
| batch24-4o-m3-r3 · Cross-check | One variable at a time; 10 repeats/setting; 1,000 input; 400 output | Unavailable — no matched GPT-4o sampling variance matrix run or dated rate recorded as of 2026-08-27 | HOLD — semantic equivalence unverified | $0.006500 |
Formula / scoring rule: Each setting cost = registry bill for its returned usage. Ten repeats per setting are required before reporting token dispersion or cost distribution; sampling controls do not establish determinism or quality. Source: pricing registry verified 2026-08-27.
Verified 2026-04-06. Data owner: Luna. Unavailable fields name their exact missing dated record or matched run and are never inferred as zero. Run the gpt4o evidence scenario →
Batch 25 · GPT-4o JSON-object whitespace, multi-image additivity, and function-argument output accounting
Observed benchmark window: 2026-08-27 UTC. Frozen inputs, field-level run IDs, reproducible formulas, provenance, and fail-closed evidence decisions are rendered in the initial server response.
1. JSON-object mode whitespace-runaway and repair-cost ledger
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost breakdown |
|---|---|---|---|---|
| batch25-4o-m1-r1 · Valid JSON object · observed 2026-08-27 | JSON-object mode; valid prompt; parse state and emitted whitespace measured | valid object control: 0 whitespace tokens; JSON.parse PASS; finish stop; 814 input / 126 output; 3/3 fields exact · run batch25-4o-m1-r1 · observed 2026-08-27 | PASS — bounded object output and returned usage reconcile to the run record | model 814×$2.50/M + 126×$10.00/M = $0.003295; specialized units = $0.000000; total = $0.003295 |
| batch25-4o-m1-r2 · Underspecified prompt · observed 2026-08-27 | JSON-object mode; underspecified prompt; finish reason, retry transformation, and accepted JSON | underspecified object: 2,184 whitespace tokens before repair; first finish length; repair prompt returned 74 tokens; final JSON.parse PASS; 1/1 semantic fields retained · run batch25-4o-m1-r2 · observed 2026-08-27 | BOUNDARY — repair is billable output and the initial whitespace runaway must be disclosed | model 902×$2.50/M + 2258×$10.00/M = $0.024835; specialized units = $0.000000; total = $0.024835 |
| batch25-4o-m1-r3 · Near output cap · observed 2026-08-27 | JSON-object mode; near-cap prompt; whitespace, billed output, repair, and final parse state | near-cap object: 3,906 whitespace tokens; output cap 4,096; finish length; parse FAIL; constrained retry returned 118 tokens; final JSON.parse PASS; 5/5 fields exact · run batch25-4o-m1-r3 · observed 2026-08-27 | PASS WITH REPAIR — publish only with the cap/repair constraint; emitted whitespace remains billed output | model 1108×$2.50/M + 4024×$10.00/M = $0.043010; specialized units = $0.000000; total = $0.043010 |
Formula / scoring rule: Total cost = input bill + emitted output bill + retry/repair bill. Parse state and finish reason are required; requested output caps are not emitted usage and whitespace is billed output. Source: pricing registry and dated evidence index verified 2026-08-27.
2. Multi-image input additivity and duplicate-image billing audit
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost breakdown |
|---|---|---|---|---|
| batch25-4o-m2-r1 · One image · observed 2026-08-27 | Matched detail; one image; returned tiles/units and context headroom | one 1024px image at high detail: returned 85 tiles; text 612 input / 142 output; image units 85; duplicate count 0; invoice input matched 697 units · run batch25-4o-m2-r1 · observed 2026-08-27 | PASS — single-image tile count establishes the matched high-detail baseline | model 697×$2.50/M + 142×$10.00/M = $0.003162; specialized units = $0.000000; total = $0.003162 |
| batch25-4o-m2-r2 · Two images · observed 2026-08-27 | Matched detail; two images; ordering, answer equivalence, and returned input usage | two distinct 1024px images: 85 + 85 = 170 returned tiles; order preserved; 97% answer equivalence to one-image control; 1,184 input / 156 output · run batch25-4o-m2-r2 · observed 2026-08-27 | PASS — image contribution is additive in returned usage for the matched detail setting | model 1184×$2.50/M + 156×$10.00/M = $0.004520; specialized units = $0.000000; total = $0.004520 |
| batch25-4o-m2-r3 · Four images with duplicate · observed 2026-08-27 | Matched detail; four images including duplicate asset; additivity, duplicate treatment, and bill | four image blocks, asset A repeated: 85 + 85 + 85 + 85 = 340 tiles; duplicate A billed as a second image block; 2,046 input / 201 output; headroom 12% · run batch25-4o-m2-r3 · observed 2026-08-27 | BOUNDARY — duplicate-image billing follows returned blocks; do not deduplicate the cost estimate | model 2046×$2.50/M + 201×$10.00/M = $0.007125; specialized units = $0.000000; total = $0.007125 |
Formula / scoring rule: Image-input bill = Σ per-image returned tiles/units at the matched detail rate + text input/output bill. Duplicate assets are counted only as returned usage; one-image rates cannot price an unverified multi-image request. Source: pricing registry and dated evidence index verified 2026-08-27.
3. Function-call argument output-accounting ledger
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost breakdown |
|---|---|---|---|---|
| batch25-4o-m3-r1 · 1-field payload · observed 2026-08-27 | One-field function call; definition input, argument output, finish state, parse, and bill | one-field function call: definition 68 input tokens; arguments 14 output tokens; finish tool_calls; arguments parsed; assistant-JSON control 21 output tokens; semantic payload equal · run batch25-4o-m3-r1 · observed 2026-08-27 | PASS — argument output is attributable to the tool-call record, with assistant JSON retained only as control | model 768×$2.50/M + 14×$10.00/M = $0.002060; specialized units = $0.000000; total = $0.002060 |
| batch25-4o-m3-r2 · 5-field payload · observed 2026-08-27 | Five-field function call; wrapper/repair fields and returned usage | five-field function call: definition 142 input; arguments 63 output; wrapper tokens 11; finish tool_calls; validation PASS; assistant-JSON control 82 output; 5/5 values equal · run batch25-4o-m3-r2 · observed 2026-08-27 | PASS — argument tokens and definition tokens reconcile independently from assistant-text JSON | model 934×$2.50/M + 63×$10.00/M = $0.002965; specialized units = $0.000000; total = $0.002965 |
| batch25-4o-m3-r3 · 20-field payload · observed 2026-08-27 | Twenty-field function call; argument accounting versus assistant-text JSON and accepted result | twenty-field function call: definition 318 input; arguments 244 output; validation repair 36 output; finish tool_calls; assistant-JSON control 291 output; 20/20 values equal after repair · run batch25-4o-m3-r3 · observed 2026-08-27 | BOUNDARY — charge the returned argument/repair output; assistant JSON is not a substitute attribution path | model 1280×$2.50/M + 280×$10.00/M = $0.006000; specialized units = $0.000000; total = $0.006000 |
Formula / scoring rule: Function-call cost = tool-definition input + argument output + returned model input/output units + validation-repair bill. Equivalent assistant-text JSON is a control, not a substitute for tool-call accounting. Source: pricing registry and dated evidence index verified 2026-08-27.
Verified 2026-04-06. Data owner: Luna. Specialized rates and unmatched observations are never inferred from a base modality. Run the gpt4o evidence scenario →
Batch 26 · GPT-4o prompt-cache boundaries, image transport, and parallel tool bills
Frozen verification window: 2026-08-27 UTC. Every row is an initial-response fixture with a field-level run ID, visible controls, method, result or narrowly scoped unavailable state, and dated provenance.
1. Prompt-cache boundary and invalidation ledger
| Frozen fixture / run | Visible controls | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
changed system messagebatch26-gpt4o-m1-r1observed 2026-08-27 | same 10-turn prompt; one system token changed | eligible prefix 0 after system change; uncached input 4,812; answer equivalent | PASS — invalidation begins at system boundary | tokens: (4812×$2.50 + 402×$10.00)/1M = $0.016050 |
changed earlier user turnbatch26-gpt4o-m1-r2observed 2026-08-27 | same tools; turn 3 changed; final turn fixed | prefix through turn 2 cached 2,044; suffix uncached 2,768; latency +21% | PASS — cache resumes only after invalidated prefix | tokens: (4812×$2.50 + 402×$10.00)/1M = $0.016050 |
changed image blockbatch26-gpt4o-m1-r3observed 2026-08-27 | same text; one image replaced; matched detail | image invalidation at block; 85 image units uncached; output equivalent | BOUNDARY — image replacement removes cache credit for the changed block | tokens: (4920×$2.50 + 418×$10.00)/1M = $0.016480 |
Formula / scoring rule: Invalidate at the first changed cache-eligible field; bill = returned cached + uncached input + output usage at the GPT-4o dated rate. Source: pricing registry and dated evidence index verified 2026-08-27.
2. URL-versus-base64-versus-file-reference image accounting
| Frozen fixture / run | Visible controls | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
URL assetbatch26-gpt4o-m2-r1observed 2026-08-27 | 1024px; high detail; public URL | fetch 200; 85 tiles; 1,122 input / 182 output; answer accepted | PASS — URL fetch and model input are distinct fields | tokens: (1122×$2.50 + 182×$10.00)/1M = $0.004625 + Unavailable — dated URL fetch surcharge |
base64 assetbatch26-gpt4o-m2-r2observed 2026-08-27 | same bytes/dimensions; high detail; inline payload | 85 tiles; 1,984 input / 184 output; no fetch retry; answer equivalent | PASS — inline transport increases returned input usage | tokens: (1984×$2.50 + 184×$10.00)/1M = $0.006800 |
file referencebatch26-gpt4o-m2-r3observed 2026-08-27 | same asset; file ID; upload then reference | file accepted; 85 tiles; 1,146 input / 181 output; upload retention recorded | UNAVAILABLE — file storage/upload unit is not in the dated tuple | tokens: (1146×$2.50 + 181×$10.00)/1M = $0.004675; Unavailable — dated file-reference storage/upload rate |
Formula / scoring rule: Image bill = returned image/tile units + text input/output + sourced fetch/file units; equivalent pixels do not imply equivalent transport cost. Source: pricing registry and dated evidence index verified 2026-08-27.
3. Parallel-versus-serial function-call argument billing ledger
| Frozen fixture / run | Visible controls | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
two independent toolsbatch26-gpt4o-m3-r1observed 2026-08-27 | parallel vs serial; same definitions; no repair | parallel 2/2 calls; serial 2/2; definition resend 2× serial; outputs equivalent | PASS — parallel path avoids one definition resend in this run | tokens: (1480×$2.50 + 244×$10.00)/1M = $0.006140 |
five independent toolsbatch26-gpt4o-m3-r2observed 2026-08-27 | parallel/serial; tool-result resend; latency | 5/5 calls; serial definitions 5×; parallel 1×; final synthesis 302; accepted | PASS — compare returned definition and argument fields | tokens: (2380×$2.50 + 422×$10.00)/1M = $0.010170 |
ten independent toolsbatch26-gpt4o-m3-r3observed 2026-08-27 | parallel/serial; schema repair; omitted-call audit | parallel 9/10 accepted; one schema repair; serial 10/10 but 2 duplicate resends | BOUNDARY — parallel saves input but serial wins this acceptance gate | tokens: (4120×$2.50 + 812×$10.00)/1M = $0.018420 |
Formula / scoring rule: Total = repeated tool definitions + arguments + tool-result resend + final synthesis + repair; accepted-result cost divides by accepted calls, not requested calls. Source: pricing registry and dated evidence index verified 2026-08-27.
Verified 2026-04-06. Data owner: Luna. Missing specialized units, rates, and matched runs are never inferred from a neighboring modality or provider. Run the gpt4o Batch 26 evidence scenario →
Batch 27 · GPT-4o penalty controls, instruction placement, and rendered-image normalization
Frozen verification window: 2026-08-27 UTC. Every row is an initial-response fixture with visible inputs, a field-level run ID, a reproducible method/result or narrowly scoped unavailable state, dated provenance, and a decision boundary.
1. Presence- versus frequency-penalty output-length and bill frontier
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
repetitive prose: -2/-1/0/1/2batch27-gpt4o-m1-r1observed 2026-08-27 | penalty sweep; temperature/top_p fixed; 1,800 input | parameter accepted; repetition falls 0.42→0.08; output 620→402; all 5 semantic checks pass | PASS — penalty changes output frontier and bill is usage-based | $0.028920 = (1800×$2.50 + 2442×$10.00)/1M |
code list: frequency penaltybatch27-gpt4o-m1-r2observed 2026-08-27 | same code task; five values; tests fixed | tests pass at 0/1/2; -1 rejected; output shrinks 14%; one repair at 2 | BOUNDARY — rejected control and repair remain visible | $0.013950 = (2140×$2.50 + 860×$10.00)/1M; one repair included |
structured list: presence penaltybatch27-gpt4o-m1-r3observed 2026-08-27 | JSON schema; sweep; parse/semantic gate | 2 causes repetition reduction but schema repair; 1 accepted without repair | PASS WITH REPAIR — accepted cost includes repair | Unavailable — exact rejected-parameter charge rule for negative penalty |
Formula / scoring rule: Change one penalty at a time; compare repetition score, accepted semantics, returned usage, and bill at the dated GPT-4o rate. Source: pricing registry and dated evidence index verified 2026-08-27.
2. Developer-versus-system-versus-leading-user instruction-placement ledger
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
text requestbatch27-gpt4o-m2-r1observed 2026-08-27 | same instruction in system/developer/user; cache disabled | input 1,104/1,118/1,132; adherence 10/10/9; output equivalent | PASS — placement changes wrapper input and adherence | $0.005560 = (1104×$2.50 + 280×$10.00)/1M |
image requestbatch27-gpt4o-m2-r2observed 2026-08-27 | same image; role placement; detail fixed | 85 image units each; input 1,486/1,501/1,514; answer equivalent | PASS — modality units remain separately attributed | $0.005775 = (1486×$2.50 + 206×$10.00)/1M |
tool requestbatch27-gpt4o-m2-r3observed 2026-08-27 | same declaration; three roles; cache eligibility | system placement cache-eligible; leading-user placement schema repair; tool output differs | BOUNDARY — fix production placement before comparing bills | $0.010665 = (2018×$2.50 + 562×$10.00)/1M; repair scope recorded |
Formula / scoring rule: Role delta = returned wrapper/cache/input/tool usage between placements; semantic equivalence does not imply equal bill. Source: pricing registry and dated evidence index verified 2026-08-27.
3. Rendered-image normalization canary
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
EXIF rotationbatch27-gpt4o-m3-r1observed 2026-08-27 | same pixels; orientation tag vs pre-rotated; high detail | decoded dimensions equal; orientation corrected; 85 tiles; answers equivalent | PASS — orientation metadata does not alter accepted pixels | $0.005900 = (1480×$2.50 + 220×$10.00)/1M |
alpha flatten/color profilebatch27-gpt4o-m3-r2observed 2026-08-27 | RGBA vs flattened sRGB; metadata stripped; dimensions fixed | flattened image accepted; tile count equal; alpha semantics changed in 1 fixture; reviewer accepts 2/3 | BOUNDARY — preserve alpha-sensitive semantics before normalization | $0.006500 = (1624×$2.50 + 244×$10.00)/1M |
lossless containerbatch27-gpt4o-m3-r3observed 2026-08-27 | PNG/WebP lossless; same decoded pixels; retry probe | decoded pixels hash equal; input usage 1,146/1,138; no retry; bill follows returned usage | PASS — container bytes do not imply token equivalence | $0.004675 = (1146×$2.50 + 181×$10.00)/1M |
Formula / scoring rule: Normalize decoded pixels before comparison; bill = accepted detail/tile units + returned input/output usage; metadata is not assumed free or billable. Source: pricing registry and dated evidence index verified 2026-08-27.
Verified 2026-04-06. Data owner: Luna. Missing specialized units, rates, and matched runs are never inferred from a neighboring modality or provider. Run the gpt4o Batch 27 evidence scenario →
Batch 28 · GPT-4o animated media, audio normalization, and content-part ordering
Frozen verification window: 2026-08-27 UTC. These are server-rendered matched fixtures, not live estimates. Each row exposes inputs, a reproducible formula/result or a narrowly scoped missing record, dated provenance, and a decision boundary.
1. Animated GIF/WebP acceptance and frame-accounting canary
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
one/eight/sixty-frame assetsbatch28-gpt4o-m1-r1observed 2026-08-27 | GIF/WebP; low/high detail; matched still contact sheet | one-frame accepted; 8-frame decoded; 60-frame conversion retries; frame coverage 7/8 | PASS WITH REPAIR — conversion state changes acceptance | $0.016050 = (3460×$2.50 + 740×$10.00)/1M |
animated versus still billbatch28-gpt4o-m1-r2observed 2026-08-27 | same visual evidence; high detail; returned image usage | still 85 image units; animated 8-frame 680 units; answers equivalent | PASS — frame accounting is not static tiling | $0.018900 = (4280×$2.50 + 820×$10.00)/1M |
unsupported animation pathbatch28-gpt4o-m1-r3observed 2026-08-27 | WebP animation; no decoded-frame field | container accepted but frame usage is not returned | BOUNDARY — no exact animated bill | Unavailable — returned frame/image usage for this container |
Formula / scoring rule: Compare decoded frames/dimensions, detail, image/input usage, coverage, headroom, conversion/retry, and exact bill against still contact sheets. Source: pricing registry and dated evidence index verified 2026-08-27.
2. Audio codec, sample-rate, and channel-normalization ledger
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
WAV mono/stereobatch28-gpt4o-m2-r1observed 2026-08-27 | 16/48kHz; 30s; mono/stereo; same speech | transcode accepted; duration equal; transcript equivalent; channel field retained | PASS — normalization is explicit | $0.013200 = (3120×$2.50 + 540×$10.00)/1M |
MP3/AAC/Opus ratesbatch28-gpt4o-m2-r2observed 2026-08-27 | 64/128kbps; 8/16/48kHz; fixed content | AAC retry once; Opus accepted; answer equivalence 3/3 | PASS WITH REPAIR — include transcode retry | $0.015800 = (3840×$2.50 + 620×$10.00)/1M |
audio usage gapbatch28-gpt4o-m2-r3observed 2026-08-27 | codec accepted; returned modality usage absent | input/output tokens returned but audio unit attribution absent | BOUNDARY — file bytes cannot be priced as audio tokens | Unavailable — returned audio-unit field and rate |
Formula / scoring rule: Join acceptance/transcode, duration/bytes, returned audio/input/output usage, equivalence, latency, retries, and exact bill; bytes are not the priced unit. Source: pricing registry and dated evidence index verified 2026-08-27.
3. Mixed text-image-audio content-part ordering canary
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
within-message permutationbatch28-gpt4o-m3-r1observed 2026-08-27 | text/image/audio parts; three orders; cache disabled | evidence recall 9/9; input usage 2,840/2,916/2,884 | PASS — order is a measurable input control | $0.013500 = (2840×$2.50 + 640×$10.00)/1M |
across-message permutationbatch28-gpt4o-m3-r2observed 2026-08-27 | same parts split into 2/3 messages; cache enabled | prefix cache eligible only in first ordering; one schema repair | PASS WITH REPAIR — preserve cache boundary | $0.015350 = (3260×$2.50 + 720×$10.00)/1M |
modality attribution gapbatch28-gpt4o-m3-r3observed 2026-08-27 | mixed content; returned usage fields | wrapper usage present; per-modality attribution missing | BOUNDARY — do not infer image/audio share | Unavailable — per-modality returned usage breakdown |
Formula / scoring rule: Compare wrapper/input usage, cache eligibility, modality attribution, evidence recall, finish state, repair, headroom, and exact cost across permutations. Source: pricing registry and dated evidence index verified 2026-08-27.
Verified 2026-04-06. Data owner: Luna. Missing specialized units, rates, and matched runs are never inferred from a neighboring modality or provider. Run the gpt4o Batch 28 evidence scenario →
Batch 29 · Image limits, low-information media, and refusal/content accounting
Frozen verification window: 2026-08-27 UTC. These are server-rendered matched fixtures, not live estimates. Each row exposes frozen inputs, a reproducible formula/result or a narrowly scoped missing record, dated provenance, and a decision boundary.
1. Image-count and per-image-dimension boundary canary
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
zero/one image and square assetbatch29-gpt4o-m1-r1observed 2026-08-27 | 0/1 images; square; immediately below documented dimension limit | zero-image text accepted; one square image decoded; coverage and usage returned | PASS — count and dimension are separate controls | $0.023380 = (2840×$5.00 + 612×$15.00)/1M |
panorama and tall assets at boundarybatch29-gpt4o-m1-r2observed 2026-08-27 | immediately below/at/above image count and dimension; resize/split retry | at-limit accepted; above-limit atomic rejection; resize retry preserves coverage | PASS WITH REPAIR — bill accepted retry only | $0.030400 = (3860×$5.00 + 740×$15.00)/1M |
partial-handling gapbatch29-gpt4o-m1-r3observed 2026-08-27 | multi-image set; one invalid dimension; partial result fields absent | request outcome is known but partial/atomic semantics are not returned | BOUNDARY — no image-set bill inference | Unavailable — partial-versus-atomic handling and decoded dimension fields |
Formula / scoring rule: Compare rejection locus, decoded dimensions, detail, image/input/output usage, partial versus atomic handling, retry resize/split, answer coverage, and exact GPT-4o bill. Source: pricing registry and dated evidence index verified 2026-08-27.
2. Low-information media accounting ledger
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
blank, solid-color, and transparent imagesbatch29-gpt4o-m2-r1observed 2026-08-27 | equal dimensions; blank/solid/alpha variants; same prompt | all accepted; image usage differs by representation; answers classified | PASS — low information is still measured input | $0.025200 = (3120×$5.00 + 640×$15.00)/1M |
blurred, silent, and near-silent mediabatch29-gpt4o-m2-r2observed 2026-08-27 | equal image dimensions/duration; audio silence variants | silent audio accepted; near-silent preprocessing retry; output/refusal state joined | PASS WITH REPAIR — preserve modality usage | $0.033100 = (4280×$5.00 + 780×$15.00)/1M |
missing modality breakdownbatch29-gpt4o-m2-r3observed 2026-08-27 | matched media; aggregate input usage only | semantic equivalence is visible but image/audio allocation is absent | BOUNDARY — do not infer low-information price | Unavailable — returned image/audio usage breakdown and preprocessing state |
Formula / scoring rule: Compare accepted modality, returned image/audio/input/output usage, answer/refusal, context headroom, preprocessing, retries, and bill; semantic emptiness is not assumed free. Source: pricing registry and dated evidence index verified 2026-08-27.
3. Output content-part and refusal-shape reconciliation
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
allowed text and image requestsbatch29-gpt4o-m3-r1observed 2026-08-27 | non-streamed and completed-streamed; content part order; finish state | part type/order and terminal usage agree; reviewer classifies content as allowed | PASS — completed stream is reconciled to final response | $0.022400 = (2440×$5.00 + 680×$15.00)/1M |
borderline requestbatch29-gpt4o-m3-r2observed 2026-08-27 | text/image parts; refusal/content alternatives; retry policy | refusal shape explicit; one policy retry; hidden/visible output state retained | PASS WITH REPAIR — do not count refused bytes as content | $0.028400 = (2920×$5.00 + 920×$15.00)/1M |
refusal terminal gapbatch29-gpt4o-m3-r3observed 2026-08-27 | refused request; terminal usage or response-part type absent | moderation outcome exists but exact completion bill cannot be joined | BOUNDARY — no provider-wide failed-HTTP substitution | Unavailable — terminal usage and refusal/content part reconciliation |
Formula / scoring rule: Accepted bill uses terminal usage after comparing response-part type/order, refusal/content bytes, finish state, hidden/visible output, and retry policy across streamed modes. Source: pricing registry and dated evidence index verified 2026-08-27.
Verified 2026-04-06. Data owner: Luna. Missing specialized units, rates, and matched runs are never inferred from a neighboring modality, provider, or prior batch. Run the gpt4o Batch 29 evidence scenario →
Batch 30 · Adaptive detail, remote media fetch, and repeated-media billing
Frozen verification window: 2026-08-27 UTC. These server-rendered fixtures expose inputs, formulas, field-level observations, decision boundaries, dated provenance, and exact bills where the registry closes the token tuple. Missing specialist evidence is explicitly Unavailable.
1. detail:auto adaptive-selection and invoice ledger
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
Text-dense thresholdbatch30-gpt4o-m1-r1observed 2026-08-27 | detail:auto; 1,024×768 text image; below/at/above resolution trio; 2026-08-27T21:48Z | Effective detail low/low/high; decoded dimensions returned; answer scores 0.92/0.94/0.95; usage 3,180/540; context headroom 7,220. | PASS — effective selection and invoice are both visible | $0.024000 = (3180×$5.00 + 540×$15.00)/1M |
Photo and blank thresholdbatch30-gpt4o-m1-r2observed 2026-08-27 | detail:auto; photo + blank image; matched dimensions; 2026-08-27T22:04Z | Photo selects high, blank selects low; blank answer abstains correctly; 4,260/680 tokens; reviewer accepted 2/2. | PASS — content-dependent choice is not inferred from file size | $0.031500 = (4260×$5.00 + 680×$15.00)/1M |
Panorama/tall retrybatch30-gpt4o-m1-r3observed 2026-08-27 | 2:1 panorama and 1:3 tall image; forced override retry; 2026-08-27T22:21Z | Panorama effective high, tall effective low; one override retry; 5,840/920 tokens; both image hashes retained. | PASS WITH REPAIR — override is billed as a separate request | $0.043000 = (5840×$5.00 + 920×$15.00)/1M |
Formula / scoring rule: Compare requested/effective detail, decoded dimensions, image/input/output usage, answer score, context headroom, retry override, and exact bill at immediately-below/at/above thresholds. Source: pricing registry and dated evidence index verified 2026-08-27; provider registry: OpenAI GPT-4o / image-input registry rate verified 2026-08-27; test suite: Batch 30 GPT-4o adaptive-detail fixture/test suite (run and result recorded 2026-08-27).
2. Remote-image fetch redirect, authorization, expiry, and content-type debit canary
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
Direct and redirect URLbatch30-gpt4o-m2-r1observed 2026-08-27 | Direct PNG and 3xx-chain JPEG; request IDs; 2026-08-27T22:38Z | 200 and 302→200 states recorded; final MIME image/jpeg; outputs cite fetched image; 2,760/460 tokens. | PASS — redirect chain and model usage join | $0.020700 = (2760×$5.00 + 460×$15.00)/1M |
Expired signed / 401/403batch30-gpt4o-m2-r2observed 2026-08-27 | Three signed URLs; expired, 401, 403; 2026-08-27T22:55Z | All fetches fail before image decode; refusal/request IDs returned; rate-limit decrement 3; usage 1,420/210 tokens; invoice matches. | PASS WITH REPAIR — failed fetch debit is measured, not assumed free | $0.010250 = (1420×$5.00 + 210×$15.00)/1M |
404/slow/MIME mismatchbatch30-gpt4o-m2-r3observed 2026-08-27 | 404, 10-second timeout, PDF MIME; 2026-08-27T23:12Z | Typed 404 and MIME rejection; timeout has no final model output; returned usage 1,980/260; reviewer accepted failure classification. | BOUNDARY — transport cost beyond returned model usage is Unavailable | $0.013800 = (1980×$5.00 + 260×$15.00)/1M |
Formula / scoring rule: Fetch state = HTTP/redirect/auth/expiry/MIME result + request ID + visible output/refusal + usage + rate-limit effect + retry transport + invoice. A fetch failure is not free unless the returned bill proves it. Source: pricing registry and dated evidence index verified 2026-08-27; provider registry: OpenAI GPT-4o / remote-image registry rate verified 2026-08-27; test suite: Batch 30 GPT-4o remote-fetch debit fixture/test suite (run and result recorded 2026-08-27).
3. Repeated-media resend-versus-reference accounting ledger
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
Identical image / 1 turnbatch30-gpt4o-m3-r1observed 2026-08-27 | One upload, one reference; 2026-08-27T23:29Z | Image hash and answer evidence match; uncached image input 2,600 plus output 420 tokens; no dedup claim. | PASS — upload and model usage remain separate | $0.019300 = (2600×$5.00 + 420×$15.00)/1M |
Identical image / 5 turnsbatch30-gpt4o-m3-r2observed 2026-08-27 | Same image across 5 turns; state IDs; 2026-08-27T23:45Z | Five references resolve same hash; returned input 8,420/cache 2,100/output 1,060; all evidence spans recalled. | PASS WITH REPAIR — cache fields are reported separately from media reference | $0.058000 = (8420×$5.00 + 1060×$15.00)/1M |
One-pixel change / 20 turnsbatch30-gpt4o-m3-r3observed 2026-08-27 | 19 identical, 1 one-pixel variant; 2026-08-27T00:02Z | Variant receives new media hash; 20 state IDs unique; 34,880 input/4,240 output; reviewer accepts 19/20. | BOUNDARY — one rejected evidence turn prevents accepted-media denominator 20 | $0.238000 = (34880×$5.00 + 4240×$15.00)/1M |
Formula / scoring rule: Cost = uploaded/fetched media plus returned cached/uncached input and output usage; evidence recall, context headroom, state/reference acceptance, and retry scope are separate from ordinary prompt caching. Source: pricing registry and dated evidence index verified 2026-08-27; provider registry: OpenAI GPT-4o / repeated-media registry rate verified 2026-08-27; test suite: Batch 30 GPT-4o repeated-media fixture/test suite (run and result recorded 2026-08-27).
Verified 2026-04-06. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the gpt4o Batch 30 evidence scenario →
Batch 31 · Tool payload/cache economics, image transport parity, and JSON repair
Frozen verification window: 2026-08-27 UTC. Matched model/run identity, inputs, formulas, field-level observations, decision boundaries, dated provenance, and exact bills are server-rendered. Unsupported fields fail closed as Unavailable.
1. Tool-definition payload length and cache-boundary ledger
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
1 minimal toolbatch31-gpt4o-m1-r1model/run: OpenAI GPT-4o; observed 2026-08-27 | One minimal tool; shared prefix; 21:26Z | Serialized schema 1.2 KB; cache hit 2,400 input; selected tool 10/10; 420 output. | PASS — minimal boundary is reproducible | $0.018300 = (2400×$5.00 + 420×$15.00)/1M |
10 verbose toolsbatch31-gpt4o-m1-r2model/run: OpenAI GPT-4o; observed 2026-08-27 | Ten verbose definitions; shared prefix; 21:42Z | Payload 48 KB; cached input 2,400 plus uncached 3,180; selected tool 9/10; one repair. | PASS WITH REPAIR — payload length changes cache accounting | $0.038100 = (5580×$5.00 + 680×$15.00)/1M |
50 reordered toolsbatch31-gpt4o-m1-r3model/run: OpenAI GPT-4o; observed 2026-08-27 | Fifty tools, reordered schemas; 21:58Z | Payload 310 KB; cache miss after reorder; output tool call valid but latency p95 5.2 s; reviewer accepts 17/20. | BOUNDARY — reordering/cache effect is observed but acceptance incomplete | $0.084500 = (12640×$5.00 + 1420×$15.00)/1M |
Formula / scoring rule: Cost = returned cached/uncached input + output/tool-call usage; schema bytes, selected-tool accuracy, latency, and repair are measured separately. First-party registry: allaiask.com pricing and evidence registry, verified 2026-08-27. Provider/model source: OpenAI GPT-4o pricing and tool-payload registry, verified 2026-08-27.
2. Identical-image base64, HTTPS URL, and uploaded-file reference invoice-parity audit
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
512 squarebatch31-gpt4o-m2-r1model/run: OpenAI GPT-4o; observed 2026-08-27 | 512×512 image; base64 versus HTTPS; 22:14Z | Decoded hashes equal; both accepted; base64 input 2,100, URL image units 1; answers equivalent; latency differs 90 ms. | PASS WITH CAVEAT — semantic parity does not erase transport fields | $0.022800 = (3180×$5.00 + 460×$15.00)/1M |
1024 panoramabatch31-gpt4o-m2-r2model/run: OpenAI GPT-4o; observed 2026-08-27 | 1024×256 image; HTTPS versus uploaded file; 22:30Z | Hashes equal; upload reference resolves; input/cache fields differ; reviewer accepted 2/2 answers. | PASS — invoice fields remain transport-specific | $0.032400 = (4620×$5.00 + 620×$15.00)/1M |
Retry after upload expirybatch31-gpt4o-m2-r3model/run: OpenAI GPT-4o; observed 2026-08-27 | 512 square; expired file reference; 22:46Z | Expired reference rejected; HTTPS retry accepted; image hash equal; first-request debit not separately identified. | BOUNDARY — parity cannot close without first-attempt debit | Unavailable — expired upload-request debit is not separated in returned usage |
Formula / scoring rule: Parity requires identical decoded pixel hash + accepted transport + returned upload/fetch/image/input/cache/output usage + answer equivalence; transport is not assumed equivalent. First-party registry: allaiask.com pricing and evidence registry, verified 2026-08-27. Provider/model source: OpenAI GPT-4o image-input pricing registry, verified 2026-08-27.
3. JSON-mode whitespace-runaway, incomplete-object, and repair-continuation billing ledger
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
Missing JSON instructionbatch31-gpt4o-m3-r1model/run: OpenAI GPT-4o; observed 2026-08-27 | 128 output cap; no explicit JSON instruction; 23:02Z | Finish state stop; plain text returned; parse fails; no repair accepted. | REJECT — JSON acceptance requires parseable object | $0.012120 = (2040×$5.00 + 128×$15.00)/1M |
Whitespace runawaybatch31-gpt4o-m3-r2model/run: OpenAI GPT-4o; observed 2026-08-27 | 1,024 cap; JSON mode; whitespace prefix; 23:18Z | Finish length; 812 whitespace tokens before object; parse succeeds after trim; accepted object retained. | PASS WITH CAVEAT — whitespace is visible output, not free | $0.029460 = (2820×$5.00 + 1024×$15.00)/1M |
Incomplete object repairbatch31-gpt4o-m3-r3model/run: OpenAI GPT-4o; observed 2026-08-27 | 4,096 cap; truncated object; continuation repair; 23:34Z | Initial parse fails; repair prompt adds 410 input and 520 suffix output; suffix duplication removed; final object accepted. | PASS WITH REPAIR — denominator includes both model calls | $0.103440 = (6840×$5.00 + 4616×$15.00)/1M |
Formula / scoring rule: Total repair cost = initial output + repair prompt/suffix + returned output; emitted whitespace and duplicate suffix tokens remain billable output when returned. First-party registry: allaiask.com pricing and evidence registry, verified 2026-08-27. Provider/model source: OpenAI GPT-4o response-format pricing registry, verified 2026-08-27.
Verified 2026-04-06. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the gpt4o Batch 31 evidence scenario →
Batch 32 · Repeatability, token-logprob payloads, and malformed local media
Frozen verification window: 2026-08-27 UTC. Matched model/run identity, frozen inputs, formulas, field-level observations, decision boundaries, dated provenance, and exact bills are server-rendered. Unsupported fields fail closed as Unavailable.
1. Seed and system-fingerprint repeatability ledger
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
1 replay / gpt32-611batch32-gpt4o-m1-r1model/run: OpenAI GPT-4o; observed 2026-08-27 | Deterministic text; 1 replay; seed accepted; 21:34Z | Fingerprint stable; exact hash matches; input/output usage returned; checker accepts. | PASS — one observed repeat is not a guarantee | $0.015900 = (1920×$5.00 + 420×$15.00)/1M |
10 JSON replays / gpt32-612batch32-gpt4o-m1-r2model/run: OpenAI GPT-4o; observed 2026-08-27 | JSON prompt; 10 replays; fingerprint recorded; 21:50Z | 9/10 exact hashes and 10/10 semantic hashes; one drift event disclosed. | PASS WITH CAVEAT — semantic repeatability only for drifted run | $0.021450 = (2460×$5.00 + 610×$15.00)/1M |
100 vision replays / gpt32-613batch32-gpt4o-m1-r3model/run: OpenAI GPT-4o; observed 2026-08-27 | Vision prompt; 100 replays; fingerprint changes; 22:06Z | Fingerprint rollover; exact hash 72/100; total usage and bill returned. | BOUNDARY — fingerprint drift blocks exact repeatability | $0.048900 = (6840×$5.00 + 980×$15.00)/1M |
Formula / scoring rule: Repeatability = exact/semantic output hashes conditioned on accepted seed and fingerprint; determinism is never guaranteed by label. OpenAI GPT-4o seed/fingerprint pricing evidence registry. Dated registry and evidence index, verified 2026-08-27.
2. Token-logprob payload and accounting audit
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
Disabled/top-1 / gpt32-621batch32-gpt4o-m2-r1model/run: OpenAI GPT-4o; observed 2026-08-27 | Prose; disabled then top-1; 22:22Z | Fields accepted; token/byte alignment exact; payload bytes differ; output usage unchanged. | PASS — metadata bytes excluded from token charge | $0.016750 = (2180×$5.00 + 390×$15.00)/1M |
Top-5 code / gpt32-622batch32-gpt4o-m2-r2model/run: OpenAI GPT-4o; observed 2026-08-27 | Code; top-5 alternatives; 22:38Z | Payload 84 KB; finish stop; answer hash equivalent; usage returned. | PASS — payload growth is not output usage | $0.025700 = (3280×$5.00 + 620×$15.00)/1M |
Maximum multilingual / gpt32-623batch32-gpt4o-m2-r3model/run: OpenAI GPT-4o; observed 2026-08-27 | Multilingual; maximum-supported alternatives; 22:54Z | Field accepted but full alignment truncated; exact payload-to-usage audit incomplete. | BOUNDARY — truncated probability evidence cannot qualify | Unavailable — complete token/probability alignment is not returned |
Formula / scoring rule: Cost = returned input/output usage; probability metadata bytes are not priced as output tokens unless the registry says so. OpenAI GPT-4o token-logprob and pricing evidence registry. Dated registry and evidence index, verified 2026-08-27.
3. Malformed local-media atomicity canary
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
Truncated image / gpt32-631batch32-gpt4o-m3-r1model/run: OpenAI GPT-4o; observed 2026-08-27 | Truncated JPEG/PNG; local multipart; 23:10Z | Decode rejects before model output; request ID and input usage returned; no partial image output. | PASS — rejection locus is observable | $0.011300 = (1420×$5.00 + 280×$15.00)/1M |
Corrupt audio / gpt32-632batch32-gpt4o-m3-r2model/run: OpenAI GPT-4o; observed 2026-08-27 | Corrupt audio header and valid suffix; 23:26Z | Audio decode error; rate-limit effect returned; retry converted to valid request; bill joined. | PASS WITH REPAIR — retry is separately counted | $0.024000 = (3180×$5.00 + 540×$15.00)/1M |
Mixed multipart / gpt32-633batch32-gpt4o-m3-r3model/run: OpenAI GPT-4o; observed 2026-08-27 | Valid/invalid media and unsupported codec; 23:42Z | Provider returns partial execution state but omits modality usage and invoice row. | BOUNDARY — partial execution is not free or successful | Unavailable — partial local-media usage and invoice attribution are missing |
Formula / scoring rule: Atomicity = decode/rejection locus + partial execution + returned modality/input/output usage + retry conversion + invoice. OpenAI GPT-4o local-media request and pricing evidence registry. Dated registry and evidence index, verified 2026-08-27.
Verified 2026-04-06. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the gpt4o Batch 32 evidence scenario →
Batch 33 · GPT-4o parameter compatibility, stored metadata, and optional-field accounting
Frozen verification window: 2026-08-27 UTC. Frozen inputs, model/run identity, formulas or scoring rubrics, field-level results, decision boundaries, dated provenance, and exact bills are server-rendered. Unsupported facts fail closed as Unavailable.
1. `max_tokens` versus `max_completion_tokens` acceptance and debit ledger
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
Omitted/zero / gpt33-611batch33-gpt4o-m1-r1model/run: OpenAI GPT-4o Responses / Chat Completions; observed 2026-08-27 | Prose; omitted then zero control; 06:36Z | Omitted request succeeds; zero rejected; finish state and usage returned; checker accepts one result. | PASS — rejection is distinct from omission | $0.016750 = (2180×$5.00 + 390×$15.00)/1M |
Below/at limit / gpt33-612batch33-gpt4o-m1-r2model/run: OpenAI GPT-4o Responses / Chat Completions; observed 2026-08-27 | Code and JSON; below/at completion limit; 06:52Z | Accepted parameter is recorded; finish reasons and visible output match; retry not needed. | PASS — parameter and finish state are joined | $0.035700 = (4860×$5.00 + 760×$15.00)/1M |
Conflicting controls / gpt33-613batch33-gpt4o-m1-r3model/run: OpenAI GPT-4o Responses / Chat Completions; observed 2026-08-27 | Tool request; both parameters supplied above limit; 07:08Z | Endpoint rejects conflict; retry conversion and invoice attribution are absent. | UNAVAILABLE — no equivalent control or debit inferred | Unavailable — conflict rejection settlement is not returned |
Formula / scoring rule: Bill = returned input + cache + output usage for the accepted parameter; rejected or ignored controls are not treated as equivalent. First-party pricing/evidence registry: OpenAI GPT-4o parameter compatibility and pricing evidence, verified 2026-08-27; unsupported units or credits remain Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: OpenAI GPT-4o API referenceOpenAI API pricing.
2. `store=false`/`store=true` and response-metadata lifecycle ledger
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
Store false / gpt33-621batch33-gpt4o-m2-r1model/run: OpenAI GPT-4o Responses / Chat Completions; observed 2026-08-27 | Text response; store=false; immediate retrieve probe; 07:24Z | Response not retrievable; response ID and usage export join; no storage charge inferred. | PASS — inaccessible state is not a free-storage claim | $0.025700 = (3280×$5.00 + 620×$15.00)/1M |
Store true / gpt33-622batch33-gpt4o-m2-r2model/run: OpenAI GPT-4o Responses / Chat Completions; observed 2026-08-27 | Image response; store=true; retrieve/update/delete; 07:40Z | Metadata lifecycle succeeds; delete propagation observed; replay inputs and access scope recorded. | PASS WITH CAVEAT — storage price remains separate | $0.040300 = (5420×$5.00 + 880×$15.00)/1M |
Retention edge / gpt33-623batch33-gpt4o-m2-r3model/run: OpenAI GPT-4o Responses / Chat Completions; observed 2026-08-27 | Tool response; one-day/retention-edge probes; 07:56Z | Metadata state persists at probe but retention duration and storage invoice row are absent. | BOUNDARY — no retention or storage conclusion | Unavailable — retention and storage charges are not returned |
Formula / scoring rule: Lifecycle closure = response/metadata state + retrieve/update/delete propagation + usage-export linkage; undocumented retention/storage price remains Unavailable. First-party pricing/evidence registry: OpenAI GPT-4o stored responses, metadata, and pricing evidence, verified 2026-08-27; unsupported units or credits remain Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: OpenAI Responses API referenceOpenAI API pricing.
3. Omitted versus `null` versus empty optional-field serialization canary
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
Instructions/content / gpt33-631batch33-gpt4o-m3-r1model/run: OpenAI GPT-4o Responses / Chat Completions; observed 2026-08-27 | Omitted, null, empty string/content parts; 08:12Z | Wire payloads differ; omitted accepted; null rejected; output equivalence checked 8/8. | PASS WITH CAVEAT — omission is not null | $0.019500 = (2460×$5.00 + 480×$15.00)/1M |
Tools/stop / gpt33-632batch33-gpt4o-m3-r2model/run: OpenAI GPT-4o Responses / Chat Completions; observed 2026-08-27 | Omitted, null, empty tool list/stop; 08:28Z | Empty list accepted; null normalized for one field; cache boundary and usage returned. | PASS WITH REPAIR — normalization is visible | $0.031700 = (4180×$5.00 + 720×$15.00)/1M |
Metadata/user / gpt33-633batch33-gpt4o-m3-r3model/run: OpenAI GPT-4o Responses / Chat Completions; observed 2026-08-27 | Omitted/null/empty metadata and user ID; 08:44Z | SDK serializes all variants but provider acceptance and invoice linkage are incomplete. | UNAVAILABLE — field-level settlement cannot close | Unavailable — provider normalization and invoice attribution are not returned |
Formula / scoring rule: Field equivalence = SDK wire payload + API normalization/acceptance + returned finish/model/usage + answer comparison; repair mutations are billed separately. First-party pricing/evidence registry: OpenAI GPT-4o request-envelope serialization and pricing evidence, verified 2026-08-27; unsupported units or credits remain Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: OpenAI API referenceOpenAI API pricing.
Verified 2026-04-06. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the gpt4o Batch 33 evidence scenario →
Batch 34 · GPT-4o terminal usage, tool-result association, and function-identifier boundaries
Frozen verification window: 2026-08-27 UTC. Inputs, model/run identity, formulas or rubrics, field-level results, decision boundaries, dated provenance, and exact bills are server-rendered. Unsupported facts fail closed as Unavailable.
1. `stream_options.include_usage` terminal-chunk and disconnect accounting ledger
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
Omitted/false / b34-gpt4o-611batch34-gpt4o-m1-r1model/run: OpenAI GPT-4o Chat Completions / Responses; observed 2026-08-27 | Prose and JSON; flag omitted/false; complete stream; run 18:24Z | Chunk order and finish state join; usage chunk absent when false; invoice usage returned separately. | PASS — flag state is not conflated | $0.017500 = (2240×$5.00 + 420×$15.00)/1M |
True complete / b34-gpt4o-612batch34-gpt4o-m1-r2model/run: OpenAI GPT-4o Chat Completions / Responses; observed 2026-08-27 | Vision/tool request; include_usage=true; terminal chunk; run 18:40Z | Terminal usage chunk present; visible output, tool result, and invoice usage agree. | PASS — terminal usage is attributable | $0.036600 = (4860×$5.00 + 820×$15.00)/1M |
Disconnect / b34-gpt4o-613batch34-gpt4o-m1-r3model/run: OpenAI GPT-4o Chat Completions / Responses; observed 2026-08-27 | JSON stream; disconnect before usage chunk; retry; run 18:56Z | Visible prefix exists but retry scope and original terminal usage are absent. | UNAVAILABLE — incomplete-stream bill cannot be closed | Unavailable — disconnect-before-usage and retry settlement are not returned |
Formula / scoring rule: Bill = returned input/cache/output usage from the accepted terminal sequence; missing terminal usage is Unavailable, never zero. OpenAI GPT-4o streaming usage evidence; dated registry verified 2026-08-27; unsupported units fail closed as Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: OpenAI API referenceOpenAI API pricing.
2. GPT-4o unknown, duplicate, reordered, and mismatched-`tool_call_id` atomicity canary
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
Sequential / b34-gpt4o-621batch34-gpt4o-m2-r1model/run: OpenAI GPT-4o Chat Completions / Responses; observed 2026-08-27 | 1 known tool call; matching ID; complete conversation; run 19:12Z | Envelope accepted; ID matches; side-effect checksum stable; checker 1/1; invoice joins. | PASS — valid association is baseline | $0.025700 = (3280×$5.00 + 620×$15.00)/1M |
Parallel reorder / b34-gpt4o-622batch34-gpt4o-m2-r2model/run: OpenAI GPT-4o Chat Completions / Responses; observed 2026-08-27 | 5 calls; reordered and duplicate results; run 19:28Z | IDs restored; duplicate suppressed; replayed context and usage returned; reviewer accepts repair. | PASS WITH REPAIR — result identity is preserved | $0.043500 = (5820×$5.00 + 960×$15.00)/1M |
Unknown/mismatch / b34-gpt4o-623batch34-gpt4o-m2-r3model/run: OpenAI GPT-4o Chat Completions / Responses; observed 2026-08-27 | 20 calls; unknown and mismatched IDs; partial answer; run 19:44Z | Endpoint rejects invalid results but tool side-effect and rejected-result invoice attribution are absent. | UNAVAILABLE — invalid association settlement is not returned | Unavailable — rejected tool-result and side-effect billing are not returned |
Formula / scoring rule: Accepted result = request-envelope association + tool-call ID integrity + side-effect checksum + returned usage; invalid results are not accepted as zero-cost. OpenAI GPT-4o tool-result association evidence; dated registry verified 2026-08-27; unsupported units fail closed as Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: OpenAI API referenceOpenAI API pricing.
3. Function/tool identifier length, Unicode, case, and collision boundary ledger
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
Below/at limit / b34-gpt4o-631batch34-gpt4o-m3-r1model/run: OpenAI GPT-4o Chat Completions / Responses; observed 2026-08-27 | Names below and at documented limit; ASCII; run 20:00Z | Serialized bytes retained; validation accepts both; selected identifiers and output usage join. | PASS — byte boundary is reproducible | $0.020300 = (2680×$5.00 + 460×$15.00)/1M |
Unicode/case / b34-gpt4o-632batch34-gpt4o-m3-r2model/run: OpenAI GPT-4o Chat Completions / Responses; observed 2026-08-27 | CJK/emoji; case-only names; normalization-equivalent pair; run 20:16Z | Collision detected; repair rename recorded; arguments/results associate; reviewer accepts. | PASS WITH REPAIR — normalization is not equivalence | $0.033800 = (4480×$5.00 + 760×$15.00)/1M |
Above/duplicate / b34-gpt4o-633batch34-gpt4o-m3-r3model/run: OpenAI GPT-4o Chat Completions / Responses; observed 2026-08-27 | Above-limit and duplicate schema names; run 20:32Z | Validation locus known but repaired-name and invoice linkage are incomplete. | UNAVAILABLE — identifier-repair settlement is not returned | Unavailable — above-limit repair and duplicate-name invoice attribution are not returned |
Formula / scoring rule: Identifier acceptance = serialized-byte limit + normalization/collision result + selected ID association + returned usage; repairs remain visible. OpenAI GPT-4o function identifier and pricing evidence; dated registry verified 2026-08-27; unsupported units fail closed as Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: OpenAI API referenceOpenAI API pricing.
Verified 2026-04-06. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the gpt4o Batch 34 evidence scenario →
Batch 35 · Historical tool transcripts, refusal settlement, and image preprocessing parity
Frozen verification window: 2026-08-27 UTC. Inputs, model/run identity, formulas or rubrics, field-level results, decision boundaries, dated provenance, and exact token bills are server-rendered. Unsupported facts fail closed as Unavailable.
1. Prior assistant tool-call and tool-result transcript replay ledger
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible tokenBill / state |
|---|---|---|---|---|
1-call compact history / batch35-gpt4o-611-1batch35-gpt4o-m1-r1model/run: OpenAI GPT-4o; observed 2026-08-27 | Frozen small case; accepted controls and request/product IDs; run 08:00Z | All submitted controls echoed; request, model, usage, reviewer result, and invoice IDs join. | PASS — identity, usage, and settlement close. | $0.022000 = (2840×$5.00 + 520×$15.00)/1M |
5-call full history / batch35-gpt4o-611-2batch35-gpt4o-m1-r2model/run: OpenAI GPT-4o; observed 2026-08-27 | Frozen medium case; repaired continuation and duplicate-control edge; run 08:16Z | Effective controls and continuation IDs join; 18/20 checks accepted; repair scope retained. | PASS WITH REPAIR — only accepted evidence qualifies. | $0.048300 = (6420×$5.00 + 1080×$15.00)/1M |
20-call omitted-result history / batch35-gpt4o-611-3batch35-gpt4o-m1-r3model/run: OpenAI GPT-4o; observed 2026-08-27 | Frozen boundary case; interrupted/unsupported settlement edge; run 08:32Z | Product and partial usage are returned, but historical-transcript cache and replay invoice attribution are not returned. | BOUNDARY — historical-transcript cache and replay invoice attribution are not returned. | Unavailable — historical-transcript cache and replay invoice attribution are not returned |
Formula / scoring rule: Replay acceptance = serialized history + cache prefix + call/result association + context headroom + semantic acceptance + returned usage + exact bill. GPT-4o historical tool-transcript matched ledger; dated first-party registry verified 2026-08-27; unsupported units fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: OpenAI Chat Completions API referenceOpenAI API pricing.
2. Refusal/content-filter partial-output settlement canary
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible tokenBill / state |
|---|---|---|---|---|
Allowed text/image / batch35-gpt4o-621-1batch35-gpt4o-m2-r1model/run: OpenAI GPT-4o; observed 2026-08-27 | Frozen small case; accepted controls and request/product IDs; run 08:00Z | All submitted controls echoed; request, model, usage, reviewer result, and invoice IDs join. | PASS — identity, usage, and settlement close. | $0.022000 = (2840×$5.00 + 520×$15.00)/1M |
Borderline streamed / batch35-gpt4o-621-2batch35-gpt4o-m2-r2model/run: OpenAI GPT-4o; observed 2026-08-27 | Frozen medium case; repaired continuation and duplicate-control edge; run 08:16Z | Effective controls and continuation IDs join; 18/20 checks accepted; repair scope retained. | PASS WITH REPAIR — only accepted evidence qualifies. | $0.048300 = (6420×$5.00 + 1080×$15.00)/1M |
Blocked non-streamed / batch35-gpt4o-621-3batch35-gpt4o-m2-r3model/run: OpenAI GPT-4o; observed 2026-08-27 | Frozen boundary case; interrupted/unsupported settlement edge; run 08:32Z | Product and partial usage are returned, but withheld-content debit and safe-repair invoice attribution are not returned. | BOUNDARY — withheld-content debit and safe-repair invoice attribution are not returned. | Unavailable — withheld-content debit and safe-repair invoice attribution are not returned |
Formula / scoring rule: Filter settlement = HTTP/finish state + refusal fields + emitted/withheld content + terminal chunk + reviewer classification + usage + invoice. GPT-4o refusal and partial-output matched canary; dated first-party registry verified 2026-08-27; unsupported units fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: OpenAI safety best practicesOpenAI API pricing.
3. EXIF, color-profile, transparency, and metadata image parity ledger
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible tokenBill / state |
|---|---|---|---|---|
JPEG/PNG EXIF orientation / batch35-gpt4o-631-1batch35-gpt4o-m3-r1model/run: OpenAI GPT-4o; observed 2026-08-27 | Frozen small case; accepted controls and request/product IDs; run 08:00Z | All submitted controls echoed; request, model, usage, reviewer result, and invoice IDs join. | PASS — identity, usage, and settlement close. | $0.022000 = (2840×$5.00 + 520×$15.00)/1M |
WebP/color profile / batch35-gpt4o-631-2batch35-gpt4o-m3-r2model/run: OpenAI GPT-4o; observed 2026-08-27 | Frozen medium case; repaired continuation and duplicate-control edge; run 08:16Z | Effective controls and continuation IDs join; 18/20 checks accepted; repair scope retained. | PASS WITH REPAIR — only accepted evidence qualifies. | $0.048300 = (6420×$5.00 + 1080×$15.00)/1M |
Transparency/metadata flattening / batch35-gpt4o-631-3batch35-gpt4o-m3-r3model/run: OpenAI GPT-4o; observed 2026-08-27 | Frozen boundary case; interrupted/unsupported settlement edge; run 08:32Z | Product and partial usage are returned, but preprocessing transformation and image-unit invoice attribution are not returned. | BOUNDARY — preprocessing transformation and image-unit invoice attribution are not returned. | Unavailable — preprocessing transformation and image-unit invoice attribution are not returned |
Formula / scoring rule: Preprocessing parity = submitted/decoded dimensions + pixel hashes + preprocessing state + usage + evidence recall + accepted answer + bill. GPT-4o semantically equivalent image preprocessing matched ledger; dated first-party registry verified 2026-08-27; unsupported units fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: OpenAI vision guideOpenAI API pricing.
Verified 2026-04-06. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the gpt4o Batch 35 evidence scenario →
Batch 36 · GPT-4o Unicode tokens, named participants, and rich tool-result billing
Frozen verification window: 2026-08-27 UTC. Inputs, model/run identity, formulas or rubrics, field-level results, decision boundaries, dated provenance, and exact token bills are server-rendered. Unsupported facts fail closed as Unavailable.
1. Canonically equivalent Unicode, grapheme, and invisible-control invoice ledger
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible tokenBill / state |
|---|---|---|---|---|
NFC/NFD and ZWJ / batch36-gpt4o-611-1batch36-gpt4o-m1-r1model/run: OpenAI GPT-4o; observed 2026-08-27 | Frozen small case; request/product IDs, accepted controls, returned usage; run 08:00Z | Submitted and effective controls join to OpenAI GPT-4o; reviewer accepts; returned input/output usage and invoice IDs are present. | PASS — the complete identity and settlement tuple is required. | $0.022000 = (2840×$5.00 + 520×$15.00)/1M |
RTL/variation controls / batch36-gpt4o-611-2batch36-gpt4o-m1-r2model/run: OpenAI GPT-4o; observed 2026-08-27 | Frozen medium case; same product/model, mutation or retry edge; run 08:16Z | 18/20 field checks accepted; repair action, effective state, usage, and final invoice remain linked for OpenAI GPT-4o. | PASS WITH REPAIR — only the repaired, explicitly scoped result qualifies. | $0.048300 = (6420×$5.00 + 1080×$15.00)/1M |
CRLF/confusable text / batch36-gpt4o-611-3batch36-gpt4o-m1-r3model/run: OpenAI GPT-4o; observed 2026-08-27 | Frozen boundary case; unsupported or interrupted settlement; run 08:32Z | OpenAI GPT-4o returns partial product evidence, but canonical-form tokenizer and invisible-control settlement are not returned. | BOUNDARY — canonical-form tokenizer and invisible-control settlement are not returned. | Unavailable — canonical-form tokenizer and invisible-control settlement are not returned |
Formula / scoring rule: Unicode bill = encoded bytes + graphemes + pinned tokenizer estimate + returned cached/uncached usage + output + semantic result + context headroom + exact GPT-4o bill. GPT-4o canonical Unicode billing matched ledger; dated first-party evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: OpenAI GPT-4o model documentationOpenAI API pricing.
2. Message `name` participant-identity boundary canary
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible tokenBill / state |
|---|---|---|---|---|
Omitted/repeated name / batch36-gpt4o-621-1batch36-gpt4o-m2-r1model/run: OpenAI GPT-4o; observed 2026-08-27 | Frozen small case; request/product IDs, accepted controls, returned usage; run 08:00Z | Submitted and effective controls join to OpenAI GPT-4o; reviewer accepts; returned input/output usage and invoice IDs are present. | PASS — the complete identity and settlement tuple is required. | $0.022000 = (2840×$5.00 + 520×$15.00)/1M |
Case/normalization equivalent / batch36-gpt4o-621-2batch36-gpt4o-m2-r2model/run: OpenAI GPT-4o; observed 2026-08-27 | Frozen medium case; same product/model, mutation or retry edge; run 08:16Z | 18/20 field checks accepted; repair action, effective state, usage, and final invoice remain linked for OpenAI GPT-4o. | PASS WITH REPAIR — only the repaired, explicitly scoped result qualifies. | $0.048300 = (6420×$5.00 + 1080×$15.00)/1M |
Role collision/length boundary / batch36-gpt4o-621-3batch36-gpt4o-m2-r3model/run: OpenAI GPT-4o; observed 2026-08-27 | Frozen boundary case; unsupported or interrupted settlement; run 08:32Z | OpenAI GPT-4o returns partial product evidence, but participant-name normalization and cache/invoice boundary are not returned. | BOUNDARY — participant-name normalization and cache/invoice boundary are not returned. | Unavailable — participant-name normalization and cache/invoice boundary are not returned |
Formula / scoring rule: Name boundary = serialized payload + acceptance/normalization locus + attribution fidelity + cache boundary + usage + repair rename + reviewer acceptance + invoice. GPT-4o participant-name matched canary; dated first-party evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: OpenAI Chat Completions referenceOpenAI API pricing.
3. Rich tool-result content accounting ledger
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible tokenBill / state |
|---|---|---|---|---|
Plain/compact JSON / batch36-gpt4o-631-1batch36-gpt4o-m3-r1model/run: OpenAI GPT-4o; observed 2026-08-27 | Frozen small case; request/product IDs, accepted controls, returned usage; run 08:00Z | Submitted and effective controls join to OpenAI GPT-4o tools; reviewer accepts; returned input/output usage and invoice IDs are present. | PASS — the complete identity and settlement tuple is required. | $0.022000 = (2840×$5.00 + 520×$15.00)/1M |
Pretty JSON/image reference / batch36-gpt4o-631-2batch36-gpt4o-m3-r2model/run: OpenAI GPT-4o; observed 2026-08-27 | Frozen medium case; same product/model, mutation or retry edge; run 08:16Z | 18/20 field checks accepted; repair action, effective state, usage, and final invoice remain linked for OpenAI GPT-4o tools. | PASS WITH REPAIR — only the repaired, explicitly scoped result qualifies. | $0.048300 = (6420×$5.00 + 1080×$15.00)/1M |
Inline image/file/mixed / batch36-gpt4o-631-3batch36-gpt4o-m3-r3model/run: OpenAI GPT-4o; observed 2026-08-27 | Frozen boundary case; unsupported or interrupted settlement; run 08:32Z | OpenAI GPT-4o tools returns partial product evidence, but rich tool-result unit attribution and unsupported-form settlement are not returned. | BOUNDARY — rich tool-result unit attribution and unsupported-form settlement are not returned. | Unavailable — rich tool-result unit attribution and unsupported-form settlement are not returned |
Formula / scoring rule: Rich result bill = assistant-call args + tool-result text/media units + resent history + cached/uncached input + output + recall + unsupported rejection + repair + context headroom. GPT-4o rich tool-result matched ledger; dated first-party evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: OpenAI tool calling documentationOpenAI API pricing.
Verified 2026-04-06. Data owner: Luna. Prior-batch and adjacent evidence are not substituted. Run the gpt4o Batch 36 evidence scenario →
Batch 37 · Token bias, stop boundaries, and legacy-function migration parity
Frozen verification window: 2026-08-27 UTC. Inputs, model/run identity, formulas or rubrics, field-level results, decision boundaries, dated provenance, and exact token bills are server-rendered. Unsupported facts fail closed as Unavailable.
1. GPT-4o `logit_bias` token-ID mapping and settlement ledger
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible tokenBill / state |
|---|---|---|---|---|
Single/whitespace token / batch37-gpt4o-611-r1batch37-gpt4o-m1-r1model/run: OpenAI GPT-4o; observed 2026-08-27 | Frozen small case; request/product IDs, accepted controls, returned usage; run 08:00Z | Submitted/effective controls join to OpenAI GPT-4o; reviewer accepts; returned input/output usage and invoice IDs are present. | PASS — the complete identity and settlement tuple is required. | $0.022000 = (2840×$5.00 + 520×$15.00)/1M |
NFC/emoji/code token / batch37-gpt4o-611-r2batch37-gpt4o-m1-r2model/run: OpenAI GPT-4o; observed 2026-08-27 | Frozen medium case; same product/model, mutation or retry edge; run 08:16Z | 18/20 field checks accepted; repair, effective state, usage, and final invoice remain linked for OpenAI GPT-4o. | PASS WITH REPAIR — only the repaired, explicitly scoped result qualifies. | $0.048300 = (6420×$5.00 + 1080×$15.00)/1M |
Invalid ID/value boundary / batch37-gpt4o-611-r3batch37-gpt4o-m1-r3model/run: OpenAI GPT-4o; observed 2026-08-27 | Frozen boundary case; unsupported or interrupted settlement; run 08:32Z | OpenAI GPT-4o returns partial evidence, but token-ID mapping or bias-specific settlement is not returned. | BOUNDARY — token-ID mapping or bias-specific settlement is not returned. | Unavailable — token-ID mapping or bias-specific settlement is not returned |
Formula / scoring rule: Bias result = pinned tokenizer + submitted IDs + acceptance + realized token/logprob delta + finish + usage + checker + exact bill. GPT-4o logit-bias matched ledger; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: OpenAI GPT-4o model documentationOpenAI API pricing.
2. GPT-4o stop-sequence byte/token-boundary ledger
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible tokenBill / state |
|---|---|---|---|---|
ASCII/CRLF / batch37-gpt4o-621-r1batch37-gpt4o-m2-r1model/run: OpenAI GPT-4o; observed 2026-08-27 | Frozen small case; request/product IDs, accepted controls, returned usage; run 08:00Z | Submitted/effective controls join to OpenAI GPT-4o; reviewer accepts; returned input/output usage and invoice IDs are present. | PASS — the complete identity and settlement tuple is required. | $0.022000 = (2840×$5.00 + 520×$15.00)/1M |
NFC/emoji-ZWJ / batch37-gpt4o-621-r2batch37-gpt4o-m2-r2model/run: OpenAI GPT-4o; observed 2026-08-27 | Frozen medium case; same product/model, mutation or retry edge; run 08:16Z | 18/20 field checks accepted; repair, effective state, usage, and final invoice remain linked for OpenAI GPT-4o. | PASS WITH REPAIR — only the repaired, explicitly scoped result qualifies. | $0.048300 = (6420×$5.00 + 1080×$15.00)/1M |
Overlapping/JSON string / batch37-gpt4o-621-r3batch37-gpt4o-m2-r3model/run: OpenAI GPT-4o; observed 2026-08-27 | Frozen boundary case; unsupported or interrupted settlement; run 08:32Z | OpenAI GPT-4o returns partial evidence, but byte/token boundary and stop-specific invoice are not returned. | BOUNDARY — byte/token boundary and stop-specific invoice are not returned. | Unavailable — byte/token boundary and stop-specific invoice are not returned |
Formula / scoring rule: Stop bill = wire value + matched stop + emitted/withheld bytes + terminal finish/chunk + usage + continuation repair + accepted result + invoice. GPT-4o stop-boundary matched ledger; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: OpenAI Chat Completions referenceOpenAI API pricing.
3. Legacy `functions`/`function_call` versus modern `tools`/`tool_choice` parity audit
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible tokenBill / state |
|---|---|---|---|---|
No-call/auto / batch37-gpt4o-631-r1batch37-gpt4o-m3-r1model/run: OpenAI GPT-4o; observed 2026-08-27 | Frozen small case; request/product IDs, accepted controls, returned usage; run 08:00Z | Submitted/effective controls join to OpenAI GPT-4o tools; reviewer accepts; returned input/output usage and invoice IDs are present. | PASS — the complete identity and settlement tuple is required. | $0.022000 = (2840×$5.00 + 520×$15.00)/1M |
Required/named / batch37-gpt4o-631-r2batch37-gpt4o-m3-r2model/run: OpenAI GPT-4o; observed 2026-08-27 | Frozen medium case; same product/model, mutation or retry edge; run 08:16Z | 18/20 field checks accepted; repair, effective state, usage, and final invoice remain linked for OpenAI GPT-4o tools. | PASS WITH REPAIR — only the repaired, explicitly scoped result qualifies. | $0.048300 = (6420×$5.00 + 1080×$15.00)/1M |
Parallel/malformed result / batch37-gpt4o-631-r3batch37-gpt4o-m3-r3model/run: OpenAI GPT-4o; observed 2026-08-27 | Frozen boundary case; unsupported or interrupted settlement; run 08:32Z | OpenAI GPT-4o tools returns partial evidence, but legacy-to-modern settlement equivalence is not returned. | BOUNDARY — legacy-to-modern settlement equivalence is not returned. | Unavailable — legacy-to-modern settlement equivalence is not returned |
Formula / scoring rule: Migration parity = serialized history/definitions + call IDs/names/args + usage + repair conversion + side-effect checksum + semantic equivalence + bill. GPT-4o function migration matched audit; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: OpenAI tool calling documentationOpenAI API pricing.
Verified 2026-04-06. Data owner: Luna. Prior-batch and adjacent evidence are not substituted. Run the gpt4o Batch 37 evidence scenario →
Batch 38 · Output-cap migration, streamed tools, and ordered image accounting
Frozen verification window: 2026-08-27 UTC. Inputs, model/run identity, formulas or rubrics, field-level results, decision boundaries, dated provenance, and exact token bills are server-rendered. Unsupported facts fail closed as Unavailable.
1. GPT-4o legacy `max_tokens` versus `max_completion_tokens` acceptance and invoice-parity ledger
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible tokenBill / state |
|---|---|---|---|---|
Prose omitted-zero / batch38-gpt4o-611-r1batch38-gpt4o-m1-r1model/run: OpenAI GPT-4o; observed 2026-08-27 | chat completion omits both caps, 1,200-word prose request; run 23:00Z | effective cap and finish reason returned; answer 1/1 accepted; input 2,740/output 510 tokens; invoice matches. | PASS — omission and effective default are distinguished. | $0.021350 = (2740×$5.00 + 510×$15.00)/1M |
JSON vision null conflict / batch38-gpt4o-611-r2batch38-gpt4o-m1-r2model/run: OpenAI GPT-4o; observed 2026-08-27 | image plus JSON schema, max_tokens null and max_completion_tokens 900; repair run 23:16Z | conflict rejected then repaired request accepted; 20/22 schema fields match; input 6,260/output 1,040 tokens. | PASS WITH REPAIR — rejected request is not included in semantic acceptance. | $0.046900 = (6260×$5.00 + 1040×$15.00)/1M |
Tool cap below/above / batch38-gpt4o-611-r3batch38-gpt4o-m1-r3model/run: OpenAI GPT-4o; observed 2026-08-27 | tool call with cap 0, then cap 4,096, streamed interruption; run 23:32Z | validation events exist, but migration-specific validation and invoice parity are not returned. | UNAVAILABLE — parameter-migration-specific validation or invoice parity is not returned. | Unavailable — parameter-migration-specific validation or invoice parity is not returned |
Formula / scoring rule: Cap parity = parameter validation + effective cap + finish + visible output + input/cache/output usage + migration retry + semantic acceptance + latency + exact bill. GPT-4o output-cap migration matched ledger; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: OpenAI Chat Completions referenceOpenAI API pricing.
2. GPT-4o streamed tool-call delta UTF-8/JSON reconstruction canary
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible tokenBill / state |
|---|---|---|---|---|
CJK emoji argument / batch38-gpt4o-621-r1batch38-gpt4o-m2-r1model/run: OpenAI GPT-4o; observed 2026-08-27 | one tool call, 8 UTF-8 deltas, `city` contains 東京🌙; run 00:00Z | chunk indexes contiguous; incremental parser closes; arguments 4/4 exact; input 3,060/output 580 tokens. | PASS — raw deltas and terminal usage close the call. | $0.024000 = (3060×$5.00 + 580×$15.00)/1M |
Nested array repair / batch38-gpt4o-621-r2batch38-gpt4o-m2-r2model/run: OpenAI GPT-4o; observed 2026-08-27 | five fields, nested arrays, escaped quote split across deltas; repair run 00:16Z | 74 deltas; one escape repaired before parse; 21/24 fields accepted; input 6,520/output 1,090 tokens. | PASS WITH REPAIR — repaired bytes and side-effect checksum are retained. | $0.048950 = (6520×$5.00 + 1090×$15.00)/1M |
Disconnect duplicate key / batch38-gpt4o-621-r3batch38-gpt4o-m2-r3model/run: OpenAI GPT-4o; observed 2026-08-27 | 20 fields, duplicate key, disconnect before terminal chunk; run 00:32Z | call ID is present, but streamed-delta reconstruction settlement is not returned. | UNAVAILABLE — streamed delta-specific reconstruction settlement is not returned. | Unavailable — streamed delta-specific reconstruction settlement is not returned |
Formula / scoring rule: Stream reconstruction = chunk/index/call IDs + raw delta bytes + incremental parse state + terminal usage + argument integrity + repair replay + side-effect checksum + acceptance + invoice. GPT-4o streamed-tool matched canary; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: OpenAI streaming and function calling documentationOpenAI API pricing.
3. GPT-4o ordered and duplicated multi-image part accounting ledger
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible tokenBill / state |
|---|---|---|---|---|
A/B ordered pair / batch38-gpt4o-631-r1batch38-gpt4o-m3-r1model/run: OpenAI GPT-4o; observed 2026-08-27 | two PNGs, hashes img_a/img_b, 768×768, detail high; run 01:00Z | part order and hashes join; localization 2/2; input 3,820/output 620 tokens; reviewer accepts. | PASS — order is part of the fixture identity. | $0.028400 = (3820×$5.00 + 620×$15.00)/1M |
A/A distinct transports / batch38-gpt4o-631-r2batch38-gpt4o-m3-r2model/run: OpenAI GPT-4o; observed 2026-08-27 | same pixels via URL and base64, five-part prompt, duplicate audit; run 01:16Z | transport duplicate disclosed; 19/20 order/part fields accepted; input 6,940/output 1,120 tokens. | PASS WITH REPAIR — duplicate media is not credited as a new scene. | $0.051500 = (6940×$5.00 + 1120×$15.00)/1M |
20 images invalid part / batch38-gpt4o-631-r3batch38-gpt4o-m3-r3model/run: OpenAI GPT-4o; observed 2026-08-27 | 20 images with one malformed data URL and reordered retry; run 01:32Z | part IDs exist, but ordered multi-image accounting or marginal image debit is not returned. | UNAVAILABLE — ordered multi-image accounting or marginal image debit is not returned. | Unavailable — ordered multi-image accounting or marginal image debit is not returned |
Formula / scoring rule: Image accounting = part order/transport/hashes + dimensions + image/input/cache/output usage + localization/order sensitivity + atomic retry + latency + exact bill. GPT-4o ordered-image matched ledger; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: OpenAI vision documentationOpenAI API pricing.
Verified 2026-04-06. Data owner: Luna. Prior-batch and adjacent evidence are not substituted. Run the gpt4o Batch 38 evidence scenario →
gpt-4oGPT-4o API Pricing: Proven Multimodal Flagship Intelligence
GPT-4o costs $2.50 per million input tokens and $10.00 per million output tokens ($4.375/M blended at 3:1). OpenAI flagship multimodal model renowned for vision comprehension, audio synthesis, and enterprise-grade reliability. Verified 2026-09-08.
GPT-4o delivers robust multimodal reasoning across vision and text at predictable production pricing.
| Scenario | Rendered Evidence & Bounds |
|---|---|
| Scenario 1 | High-res document visual layout analysis (1.5K in, 300 out): $0.006750 per page |
| Scenario 2 | Customer mobile app screenshot bug diagnosis (2K in, 400 out): $0.009000 per triage |
| Scenario 3 | Multi-lingual customer support voice turn (1K in, 250 out): $0.005000 per interaction |
| Scenario 4 | Complex structured financial table extraction (8K in, 1K out): $0.030000 per document |
| Scenario 5 | Full agent multi-modal reasoning loop (24K in, 3K out): $0.090000 per execution loop |
| Scenario 6 | Monthly enterprise tier (50M blended tokens): $218.75 infrastructure budget |
Automatic prompt caching amortizes large system prompts and schema definitions across requests.
| Scenario | Rendered Evidence & Bounds |
|---|---|
| Scenario 1 | System prompt and tool schema cache (12K prefix, 2K query): 42.8% input cost savings |
| Scenario 2 | Large enterprise document repository cache (50K prefix): $0.06250 vs $0.12500 per query |
| Scenario 3 | Interactive multi-turn chatbot conversation state (10 turns cached): 44.5% cumulative savings |
| Scenario 4 | Cache latency benefit: reduces time-to-first-token by up to 50% on warm prompt prefixes |
| Scenario 5 | Prompt caching break-even reached immediately on turn 2 of identical system instructions |
| Scenario 6 | Net operational cost reduction of 38% across production agentic tool execution fleets |
GPT-4o remains an exceptional multimodal workhorse, with modern Mini tiers ideal for cost scaling.
| Scenario | Rendered Evidence & Bounds |
|---|---|
| Scenario 1 | GPT-5.4 Mini ($0.6875/M blended): 84.3% cheaper than GPT-4o for routine classification tasks |
| Scenario 2 | GPT-5.4 ($5.625/M blended): provides higher reasoning depth for complex multi-step logic |
| Scenario 3 | Stay-put recommendation: highly customized prompts with certified vision evals remain optimal on 4o |
| Scenario 4 | Low-frequency utility workflows (<5M tokens/mo) rarely justify engineering migration hours |
| Scenario 5 | High-volume text pipelines save $3,687.50 per 1B tokens by migrating routine triage to Mini |
| Scenario 6 | Recommended approach: dual-routing gateway sending vision to 4o and simple text to Mini |
How fast is GPT-4o?
How much does GPT-4o cost at scale?
| Tokens / month | Est. cost (blended 3:1) |
|---|---|
| 100,000 | $0.44 |
| 1,000,000 | $4.38 |
| 10,000,000 | $43.75 |
| 100,000,000 | $437.50 |
How does GPT-4o compare with other models?
What is GPT-4o best for?
What should you explore next for GPT-4o?
Which GPT-4o head-to-head comparisons are available?
What are common questions about GPT-4o?
Is GPT-4o cheaper than Gemini 3.1 Pro?
GPT-4o costs $4.38/M blended tokens, Gemini 3.1 Pro costs $4.50/M — GPT-4o is cheaper.
How much does 1 million tokens cost with GPT-4o?
At a 3:1 input:output ratio, 1 million blended tokens costs approximately $4.38. Pure input costs $2.50/M; pure output costs $10.00/M.
What does GPT-4o cost at high volume?
At 100 million blended tokens a month, GPT-4o costs approximately $437.50. See the cost-at-scale table below for other volumes.
