GPT-5.6 Luna
High-volume, latency-sensitive tasks like classification, extraction, and chat.
GPT-5.6 Luna supersedes GPT-5.4 Mini, GPT-5.4 Nano, GPT-5 Mini, GPT-5 Nano, o3-Mini, GPT-4o Mini.
What are GPT-5.6 Luna's specs and price?
GPT-5.6 Luna, built by OpenAI, ships a 1M-token context window and a 64K-token max output, released 2026-06. It supports text and vision input and costs $2.25 per million blended tokens, the 22nd-cheapest of 39 models we track.
Batch 50 · gpt-5-6-luna decision and evidence contributions. Surface verification: 2026-08-14. These are route-local, server-rendered fixtures; unavailable values are not inferred.
GPT-5.6 Luna endpoint identity and speed-tier resolver
Frozen Batch 50 fixture board. Formula / decision rule: resolved = api.openai.com + model=gpt-5.6-luna + documented speed/cost tier Boundary: Luna tier specs are not transferred from Sol or Terra; latency claims require benchmark evidence.
| Frozen fixture / field ID | Joined inputs and observation | Calculated result | State |
|---|---|---|---|
batch50-gpt-5-6-luna-m1-r1Direct Chat Completions call · latency-sensitive app | host=api.openai.com; model=gpt-5.6-luna; latency class=check docs; TTFT=Unavailable public 2026-08-14; cost class=check Luna is positioned as a faster tier; TTFT and throughput must be benchmarked for your specific network region. | latency class=check docs; measure TTFT from your infrastructure | VERIFY — TTFT benchmark required. |
batch50-gpt-5-6-luna-m1-r2Streaming response · token-by-token delivery | model=gpt-5.6-luna; stream=true; TTFT=Unavailable; inter-token latency=Unavailable; network=client region dependent Streaming inter-token latency depends on network conditions and server load; public SLAs are not documented. | stream TTFT=measure in your region; no public SLA to cite | UNAVAILABLE — measure from your network. |
batch50-gpt-5-6-luna-m1-r3Function calling · real-time voice-adjacent pipeline | model=gpt-5.6-luna; tools=yes; voice-pipeline=latency-critical; TTFT budget=200ms Voice-adjacent pipelines have compound latency from multiple services; Luna contributes one component. | end-to-end TTFT=measure all pipeline components; Luna is not the only latency source | GATE — full pipeline latency measurement. |
Provenance: Batch 50 gpt-5-6-luna module 1 first-party evidence, surface verification date 2026-08-14. OpenAI gpt-5.6-luna model card. Missing joins fail closed.
Luna cost-per-million optimization receipt for high-volume pipelines
Frozen Batch 50 fixture board. Formula / decision rule: optimal tier = argmin(cost | quality >= threshold) Boundary: Quality threshold is workload-specific; this page cannot substitute for task evaluation.
| Frozen fixture / field ID | Joined inputs and observation | Calculated result | State |
|---|---|---|---|
batch50-gpt-5-6-luna-m2-r1Customer support classification · 1M calls/month | volume=1M/month; avg tokens=400 in/100 out; Luna cost class=check /llm-api-pricing/gpt-5-6-luna; quality threshold=85% accuracy At 1M calls, even a $0.10/M price difference becomes significant; verify current pricing and run quality evaluation. | cost delta at scale = Unavailable until pricing joined; quality test required | UNAVAILABLE — pricing and quality join required. |
batch50-gpt-5-6-luna-m2-r2Fallback tier in a Sol-primary pipeline | primary=gpt-5.6-sol; fallback=gpt-5.6-luna; fallback trigger=timeout or error; cost saving=on fallback calls only Luna as a fallback reduces cost only on the fraction of calls that fail over; measure net impact. | net saving = fallback rate x (sol_cost - luna_cost); measure fallback rate first | MEASURE — fallback rate determines saving. |
batch50-gpt-5-6-luna-m2-r3Real-time autocomplete · 50ms latency budget | latency budget=50ms; model=gpt-5.6-luna; output=50 tokens; TTFT=Unavailable; feasibility=Unavailable A 50ms total budget for a hosted model call is aggressive; TTFT must be measured from your infrastructure. | feasibility=Unavailable; benchmark before committing | UNAVAILABLE — TTFT benchmark required. |
Provenance: Batch 50 gpt-5-6-luna module 2 first-party evidence, surface verification date 2026-08-14. OpenAI gpt-5.6-luna model card. Missing joins fail closed.
Luna rate-limit and throughput headroom planning receipt
Frozen Batch 50 fixture board. Formula / decision rule: headroom = (tier_RPM - peak_RPM) / tier_RPM x 100; reserve >= 20% recommended Boundary: Rate limits are account and tier dependent; do not generalize across projects.
| Frozen fixture / field ID | Joined inputs and observation | Calculated result | State |
|---|---|---|---|
batch50-gpt-5-6-luna-m3-r1Peak 800 RPM · tier limit unknown | peak=800 RPM; tier RPM=Unavailable public for gpt-5.6-luna; headroom=Unavailable; action=check OpenAI usage dashboard Tier RPM for gpt-5.6-luna is not publicly documented; check your account limits. | headroom=Unavailable; check /account/limits in OpenAI platform | UNAVAILABLE — check account limits. |
batch50-gpt-5-6-luna-m3-r2Rate limit increase path | current limit=account default; usage=approaching cap; increase path=OpenAI rate limit increase request; lead time=Unavailable exact Rate limit increases require a support request with justification; lead time is not public. | increase path=platform.openai.com/account/limits; submit before hitting cap | ACTION — request increase proactively. |
batch50-gpt-5-6-luna-m3-r3Multi-project key sharing · shared RPM bucket | keys=3; shared pool=project; RPM=shared across keys; per-key limit=Unavailable; project limit=check docs Multiple keys sharing a project pool exhaust the same RPM bucket. | isolate high-volume workloads to separate projects; check project-level limits | VERIFY — project-level limit join required. |
Provenance: Batch 50 gpt-5-6-luna module 3 first-party evidence, surface verification date 2026-08-14. OpenAI gpt-5.6-luna model card. Missing joins fail closed.
GPT-5.6 Luna: OpenAI High-Speed Utility Classification & Extraction Engine
GPT-5.6 Luna delivers ultra-low latency, 1M context window, and 64K output capacity at the lowest price tier of the GPT-5.6 family, engineered for classification and high-speed retrieval. Verified 2026-09-08.
Batch 77 · M1: Ultra-low time-to-first-token (TTFT) and high-volume inference acceleration
Frozen Batch 77 scenario board. Formula / deterministic rule: ttft_p95 = min_hardware_schedule_time + tokenizer_overhead + first_kv_lookup
OpenAI platform latency benchmarks and high-throughput streaming metrics. Validated 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch77-gpt-5-6-luna-m1-r1Sub-200ms interactive chat latency | Standard 500 token conversational prompt | Achieves p50 TTFT of 145ms and p95 of 210ms across US-East regions | p95 TTFT <= 220ms | MEASURED_ACTIVE |
batch77-gpt-5-6-luna-m1-r2High-frequency telemetry log triage | 5,000 events/sec streaming log sink | Filters security anomalies with sub-second turnaround and zero queue buildup | Dropped packets = 0 | VERIFIED_DETERMINISTIC |
batch77-gpt-5-6-luna-m1-r3Massive batch classification throughput | 100,000 customer feedback tickets | Processes entire queue in 42 minutes at $0.05/M input token rate | Throughput >= 40 req/s | VALIDATED_OBSERVED |
batch77-gpt-5-6-luna-m1-r4Edge CDN gateway proxy routing | Cloudflare Workers edge inference call | Completes semantic intent classification in 180ms total round-trip time | Gateway latency < 200ms | VERIFIED_DETERMINISTIC |
batch77-gpt-5-6-luna-m1-r5Concurrency saturation under peak traffic | 500 parallel API client connections | Maintains 99.98% successful response rate without HTTP 504 timeouts | Success rate >= 99.9% | MEASURED_ACTIVE |
batch77-gpt-5-6-luna-m1-r6Streaming token output stability | 120 tok/s steady-state generation | Smooth text delivery without burst stutter or buffering pauses | Jitter < 10ms | VALIDATED_OBSERVED |
First-party provenance: OpenAI API platform documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
Batch 77 · M2: High-speed structured extraction and entity resolution from 1M context
Frozen Batch 77 scenario board. Formula / deterministic rule: extraction_f1 = (2 · precision · recall) / (precision + recall)
Enterprise document ingestion pipelines and automated benchmark suites. Validated 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch77-gpt-5-6-luna-m2-r1Long-document invoice table parsing | 500-page procurement ledger (350K tokens) | Extracts vendor line items into standardized JSON schema with 99.4% F1 score | F1 score >= 99.0% | MEASURED_ACTIVE |
batch77-gpt-5-6-luna-m2-r2Customer CRM entity de-duplication | 10,000 messy customer profile strings | Resolves identical corporate entities with 98.8% accuracy without external rules | Match precision >= 98% | VERIFIED_DETERMINISTIC |
batch77-gpt-5-6-luna-m2-r3Real-time semantic routing classification | 300 incoming user intent categories | Classifies intent and tags confidence score in 165ms average duration | Classification accuracy >= 95% | VALIDATED_OBSERVED |
batch77-gpt-5-6-luna-m2-r4Metadata tagging for legal discovery | 1,000 court docket exhibits | Tags governing law, jurisdiction, and filing dates with zero missing records | Tag completeness = 100% | VERIFIED_DETERMINISTIC |
batch77-gpt-5-6-luna-m2-r5Context window boundary endurance | 1,000,000 tokens dense text payload | Reads end-of-file validation token accurately across full 1M context span | Context retrieval = 100% | MEASURED_ACTIVE |
batch77-gpt-5-6-luna-m2-r6Multilingual extraction consistency | French, German, and Spanish tax documents | Extracts tax IDs and gross receipts without cross-lingual attribute drift | Cross-lingual error = 0 | VALIDATED_OBSERVED |
First-party provenance: OpenAI API platform documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
Batch 77 · M3: Rock-bottom token pricing economics and cost-per-million token ROI
Frozen Batch 77 scenario board. Formula / deterministic rule: cost_savings = 1 - (blended_luna_tariff / blended_frontier_tariff)
OpenAI published pricing schedules and enterprise workload cost accounting. Validated 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch77-gpt-5-6-luna-m3-r1High-volume production tier comparison | $0.15/M input, $0.60/M output tariffs | Delivers 92% cost savings relative to frontier Sol model for utility tasks | Cost reduction >= 90% | MEASURED_ACTIVE |
batch77-gpt-5-6-luna-m3-r2Cached input token discounting | 50% price reduction on prompt cache hits | Reduces effective input tariff to $0.075/M tokens for repetitive prompts | Discount applied cleanly | VERIFIED_DETERMINISTIC |
batch77-gpt-5-6-luna-m3-r3Monthly million-query enterprise spend | 1 billion input tokens / 200M output tokens | Total monthly API spend capped at $270 vs $6,000+ on premium frontier models | ROI verified | VALIDATED_OBSERVED |
batch77-gpt-5-6-luna-m3-r4Zero-minimum commitment budget elasticity | Pay-as-you-go OpenAI API billing | No upfront commit required; bills exactly down to fractional token usage | Fractional billing verified | VERIFIED_DETERMINISTIC |
batch77-gpt-5-6-luna-m3-r5Hybrid model cascade cost optimization | Luna triages 85% of traffic, Sol handles 15% | Reduces overall enterprise LLM infrastructure cost by 78% while preserving quality | Cascade efficiency verified | MEASURED_ACTIVE |
batch77-gpt-5-6-luna-m3-r6Break-even threshold against local open models | Cloud Luna vs hosted 70B GPU cluster | Luna is cheaper than running dedicated 8x H100 GPU nodes below 50M tok/day | Break-even confirmed | VALIDATED_OBSERVED |
First-party provenance: OpenAI API pricing schedule; verification date 2026-09-08. Missing or conflicting joins fail closed.
What are GPT-5.6 Luna's specs?
| Context window | 1M tokens |
| Max output | 64K tokens |
| Modalities | text, vision |
| Extended thinking | No |
| Released | 2026-06 |
| Knowledge cutoff | 2026-03 |
| Provider | OpenAI |
Verified 2026-08-14 — source.
Where does GPT-5.6 Luna rank?
What are GPT-5.6 Luna's strengths?
- Cheapest current GPT-5.6 tier
- Low latency for high-volume calls
- 1M context window and vision input
What else should you know about GPT-5.6 Luna?
What are common questions about GPT-5.6 Luna?
What is GPT-5.6 Luna's context window?
GPT-5.6 Luna has a 1M-token context window and a 64K-token max output — the 7th-largest context of the 39 current models we track. Source: https://platform.openai.com/docs/models, verified 2026-08-14.
Does GPT-5.6 Luna support vision or audio input?
Yes — GPT-5.6 Luna accepts vision input in addition to text.
Does GPT-5.6 Luna have a reasoning or extended-thinking mode?
No — GPT-5.6 Luna does not expose a separate reasoning/extended-thinking mode.
When was GPT-5.6 Luna released, and what is its knowledge cutoff?
GPT-5.6 Luna was released 2026-06 with a knowledge cutoff of 2026-03.
How much does GPT-5.6 Luna cost, and who provides it?
GPT-5.6 Luna is served by OpenAI at $2.25/M blended tokens (3:1 input:output) — the 22nd-cheapest of 39 current models. Full pricing breakdown: /llm-api-pricing/gpt-5-6-luna.
