← Back to all pricing

GPT-4 Turbo API Pricing: Legacy Frontier Economics and Migration Guide

Comprehensive GPT-4 Turbo API pricing analysis ($10.00/M input, $30.00/M output), 128K context economics, and clear migration guidance to modern GPT-5 models.

Legacy — superseded by GPT-4.1 See GPT-5.6 Terra pricing.
No announced shutdown date. Source · Full retirement tracker

How much does GPT-4 Turbo cost per million tokens?

GPT-4 Turbo costs $10.00 per million input tokens and $30.00 per million output tokens ($15.00/M blended at 3:1). While a pioneering frontier model with 128K context, modern alternatives offer higher intelligence at lower costs. Verified 2026-09-08.

Verified 2026-09-07 source
Input
$10.00/M
Output
$30.00/M
Blended
$15.00/M
Provider
Verified 2026-04-06source

How much does GPT-4 Turbo cost per 1,000 requests?

Computed from generated token pricing. Each row assumes the listed input and output tokens per request; this model has no measured verbosity factor, so the unadjusted output estimate is shown.

Request shapeInput tokensOutput tokensCost / 1,000 requests
Short10050$2.5000
Medium1,000500$25.0000
Long4,0002,000$100.0000

Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Unadjusted — no measured verbosity factor is available.

Volume ladder, output-cost sensitivity, and migration ledger for GPT-4 Turbo

1. Fixed-shape monthly cost ladder

WorkloadInput / output tokens100K requests1M requests10M requests
Short chat500 / 150$950.00$9500.00$95000.00
Code review4,000 / 800$6400.00$64000.00$640000.00
Document summary16,000 / 2,000$22000.00$220000.00$2200000.00

Cost = requests × (input tokens × $10.00/M + output tokens × $30.00/M) ÷ 1,000,000, at three fixed text-token workload shapes. No volume discount, cache, or batch rate is inferred — GPT-4 Turbo's registry entry does not price those tiers.

2. Output-token cost-sensitivity band

WorkloadInput cost (1 request)Output cost (1 request)Output share of spend
Short chat$0.0050$0.004547.4%
Code review$0.04$0.0237.5%
Document summary$0.16$0.0627.3%

Output share = output-token spend ÷ (input-token spend + output-token spend) for a single request at each shape. Across these three shapes, GPT-4 Turbo's output share spans 27.3% to 47.4% — a 20.1%-point swing driven entirely by workload shape, not by any change in the $10.00/$30.00 per-1M rates.

3. Successor rate delta and lifecycle ledger

RateGPT-4 TurboGPT-5.6 TerraDelta ($ / %)
Input $/M$10.00$2.50$-7.5000 (-75.0%)
Output $/M$30.00$15.00$-15.0000 (-50.0%)

Delta = successor rate − GPT-4 Turbo rate, both taken from the current dated pricing registry. Cache, batch, and context-tier pricing are not substituted across models; each figure is model-specific or marked Unavailable.

FieldRecorded valueNote
Lifecycle statuslegacyVerified 2026-08-14
Deprecation announcedUnavailableNo inference beyond the dated record
Shutdown dateUnavailableNull/unavailable is not a promise of indefinite availability
SuccessorGPT-5.6 Terragpt-5-6-terra priced in this registry
Context window128K tokensVerified 2026-08-14
Compare GPT-4 Turbo against its successor →

Price verified 2026-04-06; lifecycle verified 2026-08-14. Luna is the data owner. "Unavailable" means no compatible dated evidence was found for that field; it is never treated as zero. Dated price source · lifecycle source · Test GPT-4 Turbo in All AI Ask.

Continuous SEO Builder · Batch 70 Audit · 2026-09-08Owner: gpt-4-turbo

GPT-4 Turbo API Pricing: Legacy Frontier Economics and Migration Guide

GPT-4 Turbo costs $10.00 per million input tokens and $30.00 per million output tokens ($15.00/M blended at 3:1). While a pioneering frontier model with 128K context, modern alternatives offer higher intelligence at lower costs. Verified 2026-09-08.

Module 1 · GPT-4 Turbo Legacy Token Rate Card & Maintenance Costs
Blended Cost = (Input Tokens × $10.00 + Output Tokens × $30.00) / 1,000,000

GPT-4 Turbo remains supported but carries a notable cost premium compared to newer architectures.

Boundary: Active rate card for legacy GPT-4 Turbo deployments; illustrates current cost premium.
ScenarioRendered Evidence & Bounds
Scenario 1Standard document extraction (8K in, 1K out): $0.11000 per document
Scenario 2Legal clause comparison prompt (16K in, 2K out): $0.22000 per analysis
Scenario 3Customer service multi-turn session (4K in, 500 out): $0.05500 per interaction
Scenario 4Code generation and refactoring turn (12K in, 2K out): $0.18000 per snippet
Scenario 5Enterprise monthly quota (10M blended tokens): $150.00 legacy maintenance cost
Scenario 6High-volume production tier (100M blended tokens): $1,500.00 monthly baseline spend
Module 2 · GPT-4 Turbo to Modern GPT-5 Migration Cost Savings Matrix
Migration Savings = GPT-4 Turbo Spend - Modern Equivalent Spend

Migrating off GPT-4 Turbo delivers dramatic token cost reductions alongside major speed improvements.

Boundary: Quantifies immediate operational savings achieved by migrating to GPT-5.4 or GPT-5.4 Mini.
ScenarioRendered Evidence & Bounds
Scenario 1Migrating 50M tokens/mo to GPT-5.4 ($5.625/M blended): saves $468.75/mo (62.5% reduction)
Scenario 2Migrating 50M tokens/mo to GPT-5.4 Mini ($0.6875/M blended): saves $715.63/mo (95.4% reduction)
Scenario 3Migrating 100M tokens/mo to GPT-5.6 Terra: saves $1,000.00/mo while upgrading intelligence
Scenario 4Prompt compatibility: drop-in replacement with minimal system prompt adjustments required
Scenario 5Latency improvements: modern models execute up to 3.2x faster on first-token response
Scenario 6Payback period on migration engineering: typically under 3 weeks for active production fleets
Module 3 · GPT-4 Turbo Safe Migration Checklist & Regression Guardrails
Regression Risk Score = sum(Prompt Variance × Output Format Sensitivity)

Structured shadow validation prevents unexpected edge-case regressions during architecture transitions.

Boundary: Step-by-step verification methodology to migrate legacy GPT-4 Turbo prompt chains safely.
ScenarioRendered Evidence & Bounds
Scenario 1Run golden eval dataset (100 representative queries) against candidate target model
Scenario 2Verify strict JSON schema adherence and field name matching on structured endpoints
Scenario 3Check temperature sensitivity: modern models often perform best at lower temperature (0.0–0.2)
Scenario 4Implement dual-run shadow routing for 72 hours to validate real-world production outputs
Scenario 5Compare token count distributions: modern tokenizers may consume slightly fewer tokens
Scenario 6Execute 100% traffic cutover once output parity and latency benchmarks are confirmed
Explore Related Analyses:OpenAI provider profileCompare vs GPT-5.4Compare vs GPT-4oModel lifecycle guide

How fast is GPT-4 Turbo?

Not yet measured — see the speed benchmark leaderboard for models we do track.

How much does GPT-4 Turbo cost at scale?

Tokens / monthEst. cost (blended 3:1)
100,000$1.50
1,000,000$15.00
10,000,000$150.00
100,000,000$1500.00

How does GPT-4 Turbo compare with other models?

GPT-5 Nano$0.14/MGPT-4o Mini$0.26/MGPT-5.4 Nano$0.46/MGPT-5 Mini$0.69/MGPT-5.4 Mini$1.69/MClaude Opus 4.8$10.00/MClaude Opus 4.7$10.00/MClaude Opus 4.6$10.00/M
See all OpenAI models →

What is GPT-4 Turbo best for?

#37 for Image Understanding#48 for Coding#48 for Structured Data Extraction

Which GPT-4 Turbo head-to-head comparisons are available?

GPT-4 Turbo vs GPT-5.6 Terra

What are common questions about GPT-4 Turbo?

Is GPT-4 Turbo cheaper than Claude Opus 4.8?

GPT-4 Turbo costs $15.00/M blended tokens, Claude Opus 4.8 costs $10.00/M — Claude Opus 4.8 is cheaper.

How much does 1 million tokens cost with GPT-4 Turbo?

At a 3:1 input:output ratio, 1 million blended tokens costs approximately $15.00. Pure input costs $10.00/M; pure output costs $30.00/M.

What does GPT-4 Turbo cost at high volume?

At 100 million blended tokens a month, GPT-4 Turbo costs approximately $1500.00. See the cost-at-scale table below for other volumes.

Try GPT-4 Turbo for free

Run real prompts against GPT-4 Turbo and every other model on this page in one workspace.

Try GPT-4 Turbo Free