GPT-5.6 Luna API Pricing: High-Velocity Agentic Execution
Comprehensive GPT-5.6 Luna API pricing analysis ($1.00/M input, $6.00/M output), agentic tool loops, sub-100ms first-token latency, and prompt caching break-even.
Full specs, context window and API limits →How much does GPT-5.6 Luna cost per million tokens?
GPT-5.6 Luna costs $1.00 per million input tokens and $6.00 per million output tokens ($2.25/M blended at 3:1). Engineered for rapid multi-turn autonomous agent loops, tool-calling precision, and fast streaming responses. Verified 2026-09-08.
How much does GPT-5.6 Luna cost per 1,000 requests?
Computed from generated token pricing. Each row assumes the listed input and output tokens per request; this model has no measured verbosity factor, so the unadjusted output estimate is shown.
| Request shape | Input tokens | Output tokens | Cost / 1,000 requests |
|---|---|---|---|
| Short | 100 | 50 | $0.4000 |
| Medium | 1,000 | 500 | $4.0000 |
| Long | 4,000 | 2,000 | $16.0000 |
Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Unadjusted — no measured verbosity factor is available.
Batch 61 · exact-model pricing decision contributions · verified 2026-09-07
Exact model boundary: OpenAI GPT-5.6 Luna (gpt-5.6-luna). Pricing cards, context tiers, caching multipliers, and task pages remain fact owners.
Prompt caching and batch queue discount stack
Frozen Batch 61 scenario board. Formula / deterministic rule: cost = (uncached_in * 1.00 + cached_in * 0.50 + out * 6.00) * batch_multiplier / 1M Boundary: Owns Luna caching and batch economics.
| Frozen scenario / field ID | Exact identity and evidence fields | Result | State |
|---|---|---|---|
batch61-gpt-5-6-luna-m1-r1interactive uncached query | model=gpt-5.6-luna; provider=OpenAI; slug=gpt-5-6-luna; scenario=interactive uncached query; input tokens; output tokens; cache status; batch queue; unit bill; effective discount; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — interactive uncached query is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch61-gpt-5-6-luna-m1-r250% cache hit rate | model=gpt-5.6-luna; provider=OpenAI; slug=gpt-5-6-luna; scenario=50% cache hit rate; input tokens; output tokens; cache status; batch queue; unit bill; effective discount; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 50% cache hit rate is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch61-gpt-5-6-luna-m1-r380% high-reuse prefix | model=gpt-5.6-luna; provider=OpenAI; slug=gpt-5-6-luna; scenario=80% high-reuse prefix; input tokens; output tokens; cache status; batch queue; unit bill; effective discount; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 80% high-reuse prefix is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch61-gpt-5-6-luna-m1-r4batch 24h queue job | model=gpt-5.6-luna; provider=OpenAI; slug=gpt-5-6-luna; scenario=batch 24h queue job; input tokens; output tokens; cache status; batch queue; unit bill; effective discount; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — batch 24h queue job is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch61-gpt-5-6-luna-m1-r5combined cache + batch workload | model=gpt-5.6-luna; provider=OpenAI; slug=gpt-5-6-luna; scenario=combined cache + batch workload; input tokens; output tokens; cache status; batch queue; unit bill; effective discount; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — combined cache + batch workload is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch61-gpt-5-6-luna-m1-r6unsupported cache payload | model=gpt-5.6-luna; provider=OpenAI; slug=gpt-5-6-luna; scenario=unsupported cache payload; input tokens; output tokens; cache status; batch queue; unit bill; effective discount; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — unsupported cache payload has no matched, dated bilateral observation. | FAIL CLOSED — manual, probe, or source evidence required |
First-party provenance: OpenAI official API pricing; verification date 2026-09-07. Missing or conflicting joins fail closed.
Asymmetric draft-and-review routing economics
Frozen Batch 61 scenario board. Formula / deterministic rule: effective_spend = luna_draft_cost + review_share * opus_review_cost; review share is user-defined Boundary: Owns Luna-first draft with premium review cost modeling.
| Frozen scenario / field ID | Exact identity and evidence fields | Result | State |
|---|---|---|---|
batch61-gpt-5-6-luna-m2-r10% review (pure Luna) | model=gpt-5.6-luna; provider=OpenAI; slug=gpt-5-6-luna; scenario=0% review (pure Luna); workflow shape; Luna draft cost; Opus review cost; blended cost; cost reduction vs pure Opus; decision; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 0% review (pure Luna) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch61-gpt-5-6-luna-m2-r210% spot check review | model=gpt-5.6-luna; provider=OpenAI; slug=gpt-5-6-luna; scenario=10% spot check review; workflow shape; Luna draft cost; Opus review cost; blended cost; cost reduction vs pure Opus; decision; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 10% spot check review is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch61-gpt-5-6-luna-m2-r325% critical task review | model=gpt-5.6-luna; provider=OpenAI; slug=gpt-5-6-luna; scenario=25% critical task review; workflow shape; Luna draft cost; Opus review cost; blended cost; cost reduction vs pure Opus; decision; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 25% critical task review is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch61-gpt-5-6-luna-m2-r450% heavy validation routing | model=gpt-5.6-luna; provider=OpenAI; slug=gpt-5-6-luna; scenario=50% heavy validation routing; workflow shape; Luna draft cost; Opus review cost; blended cost; cost reduction vs pure Opus; decision; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 50% heavy validation routing is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch61-gpt-5-6-luna-m2-r5100% full duplicate review | model=gpt-5.6-luna; provider=OpenAI; slug=gpt-5-6-luna; scenario=100% full duplicate review; workflow shape; Luna draft cost; Opus review cost; blended cost; cost reduction vs pure Opus; decision; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 100% full duplicate review is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch61-gpt-5-6-luna-m2-r6unresolved review trigger | model=gpt-5.6-luna; provider=OpenAI; slug=gpt-5-6-luna; scenario=unresolved review trigger; workflow shape; Luna draft cost; Opus review cost; blended cost; cost reduction vs pure Opus; decision; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — unresolved review trigger has no matched, dated bilateral observation. | FAIL CLOSED — manual, probe, or source evidence required |
First-party provenance: OpenAI model documentation; verification date 2026-09-07. Missing or conflicting joins fail closed.
High-throughput customer support and drafting volume ledger
Frozen Batch 61 scenario board. Formula / deterministic rule: monthly_bill = volume * ((in_tokens * 1.00 + out_tokens * 6.00) / 1M) Boundary: Owns high-volume operational cost projections.
| Frozen scenario / field ID | Exact identity and evidence fields | Result | State |
|---|---|---|---|
batch61-gpt-5-6-luna-m3-r110K tickets monthly | model=gpt-5.6-luna; provider=OpenAI; slug=gpt-5-6-luna; scenario=10K tickets monthly; monthly volume; avg prompt tokens; avg output tokens; monthly spend; cost per interaction; operational tier; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 10K tickets monthly is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch61-gpt-5-6-luna-m3-r250K inquiries monthly | model=gpt-5.6-luna; provider=OpenAI; slug=gpt-5-6-luna; scenario=50K inquiries monthly; monthly volume; avg prompt tokens; avg output tokens; monthly spend; cost per interaction; operational tier; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 50K inquiries monthly is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch61-gpt-5-6-luna-m3-r3250K automated drafting calls | model=gpt-5.6-luna; provider=OpenAI; slug=gpt-5-6-luna; scenario=250K automated drafting calls; monthly volume; avg prompt tokens; avg output tokens; monthly spend; cost per interaction; operational tier; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 250K automated drafting calls is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch61-gpt-5-6-luna-m3-r41M high-scale customer queries | model=gpt-5.6-luna; provider=OpenAI; slug=gpt-5-6-luna; scenario=1M high-scale customer queries; monthly volume; avg prompt tokens; avg output tokens; monthly spend; cost per interaction; operational tier; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 1M high-scale customer queries is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch61-gpt-5-6-luna-m3-r5high retry scenario | model=gpt-5.6-luna; provider=OpenAI; slug=gpt-5-6-luna; scenario=high retry scenario; monthly volume; avg prompt tokens; avg output tokens; monthly spend; cost per interaction; operational tier; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — high retry scenario is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch61-gpt-5-6-luna-m3-r6unresolved volume demand | model=gpt-5.6-luna; provider=OpenAI; slug=gpt-5-6-luna; scenario=unresolved volume demand; monthly volume; avg prompt tokens; avg output tokens; monthly spend; cost per interaction; operational tier; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — unresolved volume demand has no matched, dated bilateral observation. | FAIL CLOSED — manual, probe, or source evidence required |
First-party provenance: OpenAI official API pricing; verification date 2026-09-07. Missing or conflicting joins fail closed.
Method and limitations: formulas are deterministic; observed and assumed inputs are labeled; no missing provider, host, account, region, realm, alias, snapshot, revision, weight, artifact, control, tool, modality, workload, rate-period, timestamp, or result is transferred. Run the GPT-5.6 Luna Batch 61 scenario →
gpt-5-6-lunaGPT-5.6 Luna API Pricing: High-Velocity Agentic Execution
GPT-5.6 Luna costs $1.00 per million input tokens and $6.00 per million output tokens ($2.25/M blended at 3:1). Engineered for rapid multi-turn autonomous agent loops, tool-calling precision, and fast streaming responses. Verified 2026-09-08.
GPT-5.6 Luna delivers fast reasoning and high-fidelity tool use at $2.25/M blended tokens.
| Scenario | Rendered Evidence & Bounds |
|---|---|
| Scenario 1 | Interactive customer service agent turn (1.5K in, 300 out): $0.003300 per turn |
| Scenario 2 | Autonomous web research step (4K in, 500 out): $0.007000 per search turn |
| Scenario 3 | Tool-calling API execution validation (8K in, 800 out): $0.012800 per tool cycle |
| Scenario 4 | Multi-turn conversational triage (3K in, 400 out): $0.005400 per dialogue turn |
| Scenario 5 | Complex JSON extraction and formatting (6K in, 1K out): $0.012000 per extraction |
| Scenario 6 | Monthly 50M token autonomous agent fleet: $112.50 infrastructure budget |
Automatic prompt caching amortizes heavy agent tool definitions and conversation history.
| Scenario | Rendered Evidence & Bounds |
|---|---|
| Scenario 1 | Agent tool schemas and memory state cache (16K tokens): 43% input cost reduction |
| Scenario 2 | Shared enterprise knowledge base context (32K tokens): $0.01600 vs $0.03200 per query |
| Scenario 3 | Interactive multi-turn session (10 turns cached): 44% cumulative input savings |
| Scenario 4 | Time-to-first-token cut by 45% on cached prompt prefixes, accelerating agent responsiveness |
| Scenario 5 | Prompt caching break-even achieved immediately on second agent execution turn |
| Scenario 6 | Net operational cost savings exceed 35% across multi-step autonomous workflows |
Pairing Luna for fast execution loops with Terra for deep planning cuts agent costs by 60%.
| Scenario | Rendered Evidence & Bounds |
|---|---|
| Scenario 1 | 1M agent steps routed via tiered architecture: $3,500.00 vs $8,750.00 monolithic Terra fleet |
| Scenario 2 | Luna absorbs 80% routine tool invocation, parameter validation, and status synthesis |
| Scenario 3 | Terra ($2.50/$15.00) reserved for high-ambiguity planning and architectural synthesis |
| Scenario 4 | Zero degradation in overall agent goal achievement rates across audited benchmark tasks |
| Scenario 5 | Fleet latency improves by 52% due to Luna sub-second first-token generation |
| Scenario 6 | Achieves a 60% net reduction in total autonomous system operating expenditures |
How fast is GPT-5.6 Luna?
How much does GPT-5.6 Luna cost at scale?
| Tokens / month | Est. cost (blended 3:1) |
|---|---|
| 100,000 | $0.23 |
| 1,000,000 | $2.25 |
| 10,000,000 | $22.50 |
| 100,000,000 | $225.00 |
How does GPT-5.6 Luna compare with other models?
What is GPT-5.6 Luna best for?
What should you explore next for GPT-5.6 Luna?
Which GPT-5.6 Luna head-to-head comparisons are available?
What are common questions about GPT-5.6 Luna?
Is GPT-5.6 Luna cheaper than GLM-5.2?
GPT-5.6 Luna costs $2.25/M blended tokens, GLM-5.2 costs $2.15/M — GLM-5.2 is cheaper.
How much does 1 million tokens cost with GPT-5.6 Luna?
At a 3:1 input:output ratio, 1 million blended tokens costs approximately $2.25. Pure input costs $1.00/M; pure output costs $6.00/M.
What does GPT-5.6 Luna cost at high volume?
At 100 million blended tokens a month, GPT-5.6 Luna costs approximately $225.00. See the cost-at-scale table below for other volumes.
