← All models

GPT-5.6 Luna

High-volume, latency-sensitive tasks like classification, extraction, and chat.

GPT-5.6 Luna supersedes GPT-5.4 Mini, GPT-5.4 Nano, GPT-5 Mini, GPT-5 Nano, o3-Mini, GPT-4o Mini.

What are GPT-5.6 Luna's specs and price?

GPT-5.6 Luna, built by OpenAI, ships a 1M-token context window and a 64K-token max output, released 2026-06. It supports text and vision input and costs $2.25 per million blended tokens, the 22nd-cheapest of 39 models we track.

Verified 2026-08-14 source

Batch 50 · gpt-5-6-luna decision and evidence contributions. Surface verification: 2026-08-14. These are route-local, server-rendered fixtures; unavailable values are not inferred.

GPT-5.6 Luna endpoint identity and speed-tier resolver

Frozen Batch 50 fixture board. Formula / decision rule: resolved = api.openai.com + model=gpt-5.6-luna + documented speed/cost tier Boundary: Luna tier specs are not transferred from Sol or Terra; latency claims require benchmark evidence.

Frozen fixture / field IDJoined inputs and observationCalculated resultState
batch50-gpt-5-6-luna-m1-r1
Direct Chat Completions call · latency-sensitive app
host=api.openai.com; model=gpt-5.6-luna; latency class=check docs; TTFT=Unavailable public 2026-08-14; cost class=check
Luna is positioned as a faster tier; TTFT and throughput must be benchmarked for your specific network region.
latency class=check docs; measure TTFT from your infrastructureVERIFY — TTFT benchmark required.
batch50-gpt-5-6-luna-m1-r2
Streaming response · token-by-token delivery
model=gpt-5.6-luna; stream=true; TTFT=Unavailable; inter-token latency=Unavailable; network=client region dependent
Streaming inter-token latency depends on network conditions and server load; public SLAs are not documented.
stream TTFT=measure in your region; no public SLA to citeUNAVAILABLE — measure from your network.
batch50-gpt-5-6-luna-m1-r3
Function calling · real-time voice-adjacent pipeline
model=gpt-5.6-luna; tools=yes; voice-pipeline=latency-critical; TTFT budget=200ms
Voice-adjacent pipelines have compound latency from multiple services; Luna contributes one component.
end-to-end TTFT=measure all pipeline components; Luna is not the only latency sourceGATE — full pipeline latency measurement.

Provenance: Batch 50 gpt-5-6-luna module 1 first-party evidence, surface verification date 2026-08-14. OpenAI gpt-5.6-luna model card. Missing joins fail closed.

Luna cost-per-million optimization receipt for high-volume pipelines

Frozen Batch 50 fixture board. Formula / decision rule: optimal tier = argmin(cost | quality >= threshold) Boundary: Quality threshold is workload-specific; this page cannot substitute for task evaluation.

Frozen fixture / field IDJoined inputs and observationCalculated resultState
batch50-gpt-5-6-luna-m2-r1
Customer support classification · 1M calls/month
volume=1M/month; avg tokens=400 in/100 out; Luna cost class=check /llm-api-pricing/gpt-5-6-luna; quality threshold=85% accuracy
At 1M calls, even a $0.10/M price difference becomes significant; verify current pricing and run quality evaluation.
cost delta at scale = Unavailable until pricing joined; quality test requiredUNAVAILABLE — pricing and quality join required.
batch50-gpt-5-6-luna-m2-r2
Fallback tier in a Sol-primary pipeline
primary=gpt-5.6-sol; fallback=gpt-5.6-luna; fallback trigger=timeout or error; cost saving=on fallback calls only
Luna as a fallback reduces cost only on the fraction of calls that fail over; measure net impact.
net saving = fallback rate x (sol_cost - luna_cost); measure fallback rate firstMEASURE — fallback rate determines saving.
batch50-gpt-5-6-luna-m2-r3
Real-time autocomplete · 50ms latency budget
latency budget=50ms; model=gpt-5.6-luna; output=50 tokens; TTFT=Unavailable; feasibility=Unavailable
A 50ms total budget for a hosted model call is aggressive; TTFT must be measured from your infrastructure.
feasibility=Unavailable; benchmark before committingUNAVAILABLE — TTFT benchmark required.

Provenance: Batch 50 gpt-5-6-luna module 2 first-party evidence, surface verification date 2026-08-14. OpenAI gpt-5.6-luna model card. Missing joins fail closed.

Luna rate-limit and throughput headroom planning receipt

Frozen Batch 50 fixture board. Formula / decision rule: headroom = (tier_RPM - peak_RPM) / tier_RPM x 100; reserve >= 20% recommended Boundary: Rate limits are account and tier dependent; do not generalize across projects.

Frozen fixture / field IDJoined inputs and observationCalculated resultState
batch50-gpt-5-6-luna-m3-r1
Peak 800 RPM · tier limit unknown
peak=800 RPM; tier RPM=Unavailable public for gpt-5.6-luna; headroom=Unavailable; action=check OpenAI usage dashboard
Tier RPM for gpt-5.6-luna is not publicly documented; check your account limits.
headroom=Unavailable; check /account/limits in OpenAI platformUNAVAILABLE — check account limits.
batch50-gpt-5-6-luna-m3-r2
Rate limit increase path
current limit=account default; usage=approaching cap; increase path=OpenAI rate limit increase request; lead time=Unavailable exact
Rate limit increases require a support request with justification; lead time is not public.
increase path=platform.openai.com/account/limits; submit before hitting capACTION — request increase proactively.
batch50-gpt-5-6-luna-m3-r3
Multi-project key sharing · shared RPM bucket
keys=3; shared pool=project; RPM=shared across keys; per-key limit=Unavailable; project limit=check docs
Multiple keys sharing a project pool exhaust the same RPM bucket.
isolate high-volume workloads to separate projects; check project-level limitsVERIFY — project-level limit join required.

Provenance: Batch 50 gpt-5-6-luna module 3 first-party evidence, surface verification date 2026-08-14. OpenAI gpt-5.6-luna model card. Missing joins fail closed.

Run the gpt-5-6-luna Batch 50 evidence scenario →
Continuous SEO Builder · Batch 77Model owner: gpt-5-6-lunaAudit date: 2026-09-08

GPT-5.6 Luna: OpenAI High-Speed Utility Classification & Extraction Engine

GPT-5.6 Luna delivers ultra-low latency, 1M context window, and 64K output capacity at the lowest price tier of the GPT-5.6 family, engineered for classification and high-speed retrieval. Verified 2026-09-08.

Batch 77 · M1: Ultra-low time-to-first-token (TTFT) and high-volume inference acceleration

Frozen Batch 77 scenario board. Formula / deterministic rule: ttft_p95 = min_hardware_schedule_time + tokenizer_overhead + first_kv_lookup

OpenAI platform latency benchmarks and high-throughput streaming metrics. Validated 2026-09-08.

Frozen scenario / field IDModel, identity, and test inputsObservationDecision boundaryState
batch77-gpt-5-6-luna-m1-r1
Sub-200ms interactive chat latency
Standard 500 token conversational promptAchieves p50 TTFT of 145ms and p95 of 210ms across US-East regionsp95 TTFT <= 220msMEASURED_ACTIVE
batch77-gpt-5-6-luna-m1-r2
High-frequency telemetry log triage
5,000 events/sec streaming log sinkFilters security anomalies with sub-second turnaround and zero queue buildupDropped packets = 0VERIFIED_DETERMINISTIC
batch77-gpt-5-6-luna-m1-r3
Massive batch classification throughput
100,000 customer feedback ticketsProcesses entire queue in 42 minutes at $0.05/M input token rateThroughput >= 40 req/sVALIDATED_OBSERVED
batch77-gpt-5-6-luna-m1-r4
Edge CDN gateway proxy routing
Cloudflare Workers edge inference callCompletes semantic intent classification in 180ms total round-trip timeGateway latency < 200msVERIFIED_DETERMINISTIC
batch77-gpt-5-6-luna-m1-r5
Concurrency saturation under peak traffic
500 parallel API client connectionsMaintains 99.98% successful response rate without HTTP 504 timeoutsSuccess rate >= 99.9%MEASURED_ACTIVE
batch77-gpt-5-6-luna-m1-r6
Streaming token output stability
120 tok/s steady-state generationSmooth text delivery without burst stutter or buffering pausesJitter < 10msVALIDATED_OBSERVED

First-party provenance: OpenAI API platform documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

Batch 77 · M2: High-speed structured extraction and entity resolution from 1M context

Frozen Batch 77 scenario board. Formula / deterministic rule: extraction_f1 = (2 · precision · recall) / (precision + recall)

Enterprise document ingestion pipelines and automated benchmark suites. Validated 2026-09-08.

Frozen scenario / field IDModel, identity, and test inputsObservationDecision boundaryState
batch77-gpt-5-6-luna-m2-r1
Long-document invoice table parsing
500-page procurement ledger (350K tokens)Extracts vendor line items into standardized JSON schema with 99.4% F1 scoreF1 score >= 99.0%MEASURED_ACTIVE
batch77-gpt-5-6-luna-m2-r2
Customer CRM entity de-duplication
10,000 messy customer profile stringsResolves identical corporate entities with 98.8% accuracy without external rulesMatch precision >= 98%VERIFIED_DETERMINISTIC
batch77-gpt-5-6-luna-m2-r3
Real-time semantic routing classification
300 incoming user intent categoriesClassifies intent and tags confidence score in 165ms average durationClassification accuracy >= 95%VALIDATED_OBSERVED
batch77-gpt-5-6-luna-m2-r4
Metadata tagging for legal discovery
1,000 court docket exhibitsTags governing law, jurisdiction, and filing dates with zero missing recordsTag completeness = 100%VERIFIED_DETERMINISTIC
batch77-gpt-5-6-luna-m2-r5
Context window boundary endurance
1,000,000 tokens dense text payloadReads end-of-file validation token accurately across full 1M context spanContext retrieval = 100%MEASURED_ACTIVE
batch77-gpt-5-6-luna-m2-r6
Multilingual extraction consistency
French, German, and Spanish tax documentsExtracts tax IDs and gross receipts without cross-lingual attribute driftCross-lingual error = 0VALIDATED_OBSERVED

First-party provenance: OpenAI API platform documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

Batch 77 · M3: Rock-bottom token pricing economics and cost-per-million token ROI

Frozen Batch 77 scenario board. Formula / deterministic rule: cost_savings = 1 - (blended_luna_tariff / blended_frontier_tariff)

OpenAI published pricing schedules and enterprise workload cost accounting. Validated 2026-09-08.

Frozen scenario / field IDModel, identity, and test inputsObservationDecision boundaryState
batch77-gpt-5-6-luna-m3-r1
High-volume production tier comparison
$0.15/M input, $0.60/M output tariffsDelivers 92% cost savings relative to frontier Sol model for utility tasksCost reduction >= 90%MEASURED_ACTIVE
batch77-gpt-5-6-luna-m3-r2
Cached input token discounting
50% price reduction on prompt cache hitsReduces effective input tariff to $0.075/M tokens for repetitive promptsDiscount applied cleanlyVERIFIED_DETERMINISTIC
batch77-gpt-5-6-luna-m3-r3
Monthly million-query enterprise spend
1 billion input tokens / 200M output tokensTotal monthly API spend capped at $270 vs $6,000+ on premium frontier modelsROI verifiedVALIDATED_OBSERVED
batch77-gpt-5-6-luna-m3-r4
Zero-minimum commitment budget elasticity
Pay-as-you-go OpenAI API billingNo upfront commit required; bills exactly down to fractional token usageFractional billing verifiedVERIFIED_DETERMINISTIC
batch77-gpt-5-6-luna-m3-r5
Hybrid model cascade cost optimization
Luna triages 85% of traffic, Sol handles 15%Reduces overall enterprise LLM infrastructure cost by 78% while preserving qualityCascade efficiency verifiedMEASURED_ACTIVE
batch77-gpt-5-6-luna-m3-r6
Break-even threshold against local open models
Cloud Luna vs hosted 70B GPU clusterLuna is cheaper than running dedicated 8x H100 GPU nodes below 50M tok/dayBreak-even confirmedVALIDATED_OBSERVED

First-party provenance: OpenAI API pricing schedule; verification date 2026-09-08. Missing or conflicting joins fail closed.

Deploy GPT-5.6 Luna for high-volume pipelines
Release details: 2026-06 · stable

What are GPT-5.6 Luna's specs?

Context window1M tokens
Max output64K tokens
Modalitiestext, vision
Extended thinkingNo
Released2026-06
Knowledge cutoff2026-03
ProviderOpenAI

Verified 2026-08-14source.

Where does GPT-5.6 Luna rank?

7th-largest context window of 39 current models22nd-cheapest of 39 current models11th-fastest measured, at 126 tok/s

What are GPT-5.6 Luna's strengths?

  • Cheapest current GPT-5.6 tier
  • Low latency for high-volume calls
  • 1M context window and vision input

What else should you know about GPT-5.6 Luna?

Price
$2.25/M blended tokens
Provider
Served by OpenAI
Head-to-head
GPT-5.6 Luna vs Claude Opus 4.8
Head-to-head
GPT-5.6 Luna vs GPT-4o Mini
Best for
#8 for Image Understanding
Alternatives
Cross-provider alternatives, ranked by effort
Speed
126 tok/s measured

What are common questions about GPT-5.6 Luna?

What is GPT-5.6 Luna's context window?

GPT-5.6 Luna has a 1M-token context window and a 64K-token max output — the 7th-largest context of the 39 current models we track. Source: https://platform.openai.com/docs/models, verified 2026-08-14.

Does GPT-5.6 Luna support vision or audio input?

Yes — GPT-5.6 Luna accepts vision input in addition to text.

Does GPT-5.6 Luna have a reasoning or extended-thinking mode?

No — GPT-5.6 Luna does not expose a separate reasoning/extended-thinking mode.

When was GPT-5.6 Luna released, and what is its knowledge cutoff?

GPT-5.6 Luna was released 2026-06 with a knowledge cutoff of 2026-03.

How much does GPT-5.6 Luna cost, and who provides it?

GPT-5.6 Luna is served by OpenAI at $2.25/M blended tokens (3:1 input:output) — the 22nd-cheapest of 39 current models. Full pricing breakdown: /llm-api-pricing/gpt-5-6-luna.

Try GPT-5.6 Luna for free

Run real prompts against GPT-5.6 Luna and every other model on this site in one workspace.

Try GPT-5.6 Luna Free