← Back to all pricing

o3-Mini API Pricing: High-Speed Analytical and STEM Reasoning

Comprehensive o3-Mini API pricing analysis ($1.10/M input, $4.40/M output), reasoning token overhead, STEM calculation efficiency, and coding benchmarks.

Legacy — use GPT-5.6 Luna See GPT-5.6 Luna pricing.
No announced shutdown date. Source · Full retirement tracker

How much does o3-Mini cost per million tokens?

o3-Mini costs $1.10 per million input tokens and $4.40 per million output tokens ($1.925/M blended at 3:1). Optimized for high-throughput mathematical reasoning, coding challenges, and logical analysis. Verified 2026-09-08.

Verified 2026-09-07 source
Input
$1.10/M
Output
$4.40/M
Blended
$1.93/M
Provider
Verified 2026-04-06source

How much does o3-Mini cost per 1,000 requests?

Computed from generated token pricing. Each row assumes the listed input and output tokens per request; this model has no measured verbosity factor, so the unadjusted output estimate is shown.

Request shapeInput tokensOutput tokensCost / 1,000 requests
Short10050$0.3300
Medium1,000500$3.3000
Long4,0002,000$13.2000

Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Unadjusted — no measured verbosity factor is available.

Volume ladder, output-cost sensitivity, and migration ledger for o3-Mini

1. Fixed-shape monthly cost ladder

WorkloadInput / output tokens100K requests1M requests10M requests
Short chat500 / 150$121.00$1210.00$12100.00
Code review4,000 / 800$792.00$7920.00$79200.00
Document summary16,000 / 2,000$2640.00$26400.00$264000.00

Cost = requests × (input tokens × $1.10/M + output tokens × $4.40/M) ÷ 1,000,000, at three fixed text-token workload shapes. No volume discount, cache, or batch rate is inferred — o3-Mini's registry entry does not price those tiers.

2. Output-token cost-sensitivity band

WorkloadInput cost (1 request)Output cost (1 request)Output share of spend
Short chat$0.0006$0.000754.5%
Code review$0.0044$0.003544.4%
Document summary$0.02$0.008833.3%

Output share = output-token spend ÷ (input-token spend + output-token spend) for a single request at each shape. Across these three shapes, o3-Mini's output share spans 33.3% to 54.5% — a 21.2%-point swing driven entirely by workload shape, not by any change in the $1.10/$4.40 per-1M rates.

3. Successor rate delta and lifecycle ledger

Rateo3-MiniGPT-5.6 LunaDelta ($ / %)
Input $/M$1.10$1.00$-0.1000 (-9.1%)
Output $/M$4.40$6.00$1.60 (36.4%)

Delta = successor rate − o3-Mini rate, both taken from the current dated pricing registry. Cache, batch, and context-tier pricing are not substituted across models; each figure is model-specific or marked Unavailable.

FieldRecorded valueNote
Lifecycle statuslegacyVerified 2026-08-14
Deprecation announcedUnavailableNo inference beyond the dated record
Shutdown dateUnavailableNull/unavailable is not a promise of indefinite availability
SuccessorGPT-5.6 Lunagpt-5-6-luna priced in this registry
Context windowUnavailableUnavailable — no model-specific spec sourced
Compare o3-Mini against its successor →

Price verified 2026-04-06; lifecycle verified 2026-08-14. Luna is the data owner. "Unavailable" means no compatible dated evidence was found for that field; it is never treated as zero. Dated price source · lifecycle source · Test o3-Mini in All AI Ask.

Continuous SEO Builder · Batch 70 Audit · 2026-09-08Owner: o3-mini

o3-Mini API Pricing: High-Speed Analytical and STEM Reasoning

o3-Mini costs $1.10 per million input tokens and $4.40 per million output tokens ($1.925/M blended at 3:1). Optimized for high-throughput mathematical reasoning, coding challenges, and logical analysis. Verified 2026-09-08.

Module 1 · o3-Mini Mathematical Reasoning Unit Token Economics
Blended Cost = (Input Tokens × $1.10 + Output & Reasoning Tokens × $4.40) / 1,000,000

o3-Mini delivers deep logical verification at a fraction of full-size frontier reasoning costs.

Boundary: Reasoning tokens count toward output token billing at standard $4.40/M rate.
ScenarioRendered Evidence & Bounds
Scenario 1Competitive programming LeetCode hard problem (1K in, 2.5K thought+out): $0.01210 per run
Scenario 2Calculus proof verification (2K in, 4K thought+out): $0.01980 per problem
Scenario 3Algorithmic trading strategy backtest verification (4K in, 6K thought+out): $0.03080 per run
Scenario 4Symbolic logic constraint solver (3K in, 5K thought+out): $0.02530 per problem
Scenario 5High-throughput STEM grading engine (500 in, 1.5K thought+out): $0.00715 per submission
Scenario 6Monthly educational platform run (20M tokens): $38.50 total infrastructure cost
Module 2 · o3-Mini Reasoning Effort Parameter Tuning & Spend Control
Spend Control = Adjust reasoning_effort (low / medium / high) to cap thought tokens

Configuring reasoning effort ensures deep thinking is only applied when mathematical rigor requires it.

Boundary: Evaluates trade-off between reasoning token consumption and solution correctness.
ScenarioRendered Evidence & Bounds
Scenario 1Low effort (avg 800 reasoning tokens): 65% cost reduction on routine algebraic tasks
Scenario 2Medium effort (avg 2,200 reasoning tokens): optimal balance for coding challenges
Scenario 3High effort (avg 6,500 reasoning tokens): maximum accuracy for formal mathematical proofs
Scenario 4Reasoning token budget caps prevent runaway chain-of-thought execution loops
Scenario 5Automated effort escalation: low on initial attempt, medium only on failing unit tests
Scenario 6Reduces average reasoning query spend from $0.03080 to $0.01050 across production fleets
Module 3 · o3-Mini vs Standard Model Fleet Cost-Performance Frontier
Fleet Efficiency = (o3-Mini Accuracy × Task Value) - Total Reasoning Spend

First-pass correctness on reasoning models often costs less than multi-turn retry loops with standard LLMs.

Boundary: Measures whether higher token counts on reasoning models produce better net cost efficiency.
ScenarioRendered Evidence & Bounds
Scenario 1Standard model requires 4 retries ($0.024) vs o3-Mini single-shot success ($0.019): o3-Mini wins
Scenario 2Zero hallucination on complex tax code calculations: eliminates costly manual re-audits
Scenario 3Automated SQL query generation with zero syntax errors: 99.4% first-run validation rate
Scenario 4Unit test generation: achieves 94% branch coverage on first generation attempt
Scenario 510K complex engineering queries: $192.50 total spend vs $340.00 multi-step agent loops
Scenario 6Overall fleet compute efficiency improves by 43% when delegating STEM tasks to o3-Mini
Explore Related Analyses:OpenAI provider profileCompare vs GPT-5 MiniCompare vs DeepSeek V4 FlashMath and reasoning benchmark

How fast is o3-Mini?

Not yet measured — see the speed benchmark leaderboard for models we do track.

How much does o3-Mini cost at scale?

Tokens / monthEst. cost (blended 3:1)
100,000$0.19
1,000,000$1.93
10,000,000$19.25
100,000,000$192.50

How does o3-Mini compare with other models?

GPT-5 Nano$0.14/MGPT-4o Mini$0.26/MGPT-5.4 Nano$0.46/MGPT-5 Mini$0.69/MGPT-5.4 Mini$1.69/MDeepSeek V4 Pro$1.98/MClaude Haiku 4.5$2.00/MMuse Spark 1.3$2.00/M
See all OpenAI models →

What are common questions about o3-Mini?

Is o3-Mini cheaper than DeepSeek V4 Pro?

o3-Mini costs $1.93/M blended tokens, DeepSeek V4 Pro costs $1.98/M — o3-Mini is cheaper.

How much does 1 million tokens cost with o3-Mini?

At a 3:1 input:output ratio, 1 million blended tokens costs approximately $1.93. Pure input costs $1.10/M; pure output costs $4.40/M.

What does o3-Mini cost at high volume?

At 100 million blended tokens a month, o3-Mini costs approximately $192.50. See the cost-at-scale table below for other volumes.

Try o3-Mini for free

Run real prompts against o3-Mini and every other model on this page in one workspace.

Try o3-Mini Free