LLM API Pricing Comparison — 68 Models, Updated April 2026

We compare standard API rates across 68 models, from Amazon Nova Micro at $0.061 per million blended tokens to the most expensive model at GPT-5.4 Pro at $67.50 per million — a 1102x spread. Every price below is sourced directly from the provider's official pricing page and dated with the last verification. Use the calculator to estimate your own monthly cost across all 68 models.

Quick answer

How should I compare LLM API prices?

Compare input, cached-input, output, blended workload cost, context, and freshness in one table. For the single lowest-cost recommendation, use the cheapest LLM API ranking; this page remains the source comparison for the underlying prices.

The "Your cost" column updates live below.
ModelProvider$ / M input$ / M outputBlendedYour cost / mo
Amazon Nova MicroAmazon$0.03$0.14$0.06$0.07
Amazon Nova LiteAmazon$0.06$0.24$0.11$0.12
Muse Spark 1.3 ContributorMeta$0.10$0.20$0.13$0.15
GPT-OSS 20BGroq$0.07$0.30$0.13$0.15
Ministral 8BMistral$0.15$0.15$0.15$0.19
GPT-OSS 120BGroq$0.15$0.60$0.26$0.30
Mistral Small 3.1Mistral$0.15$0.60$0.26$0.30
CodestralMistral$0.30$0.90$0.45$0.53
GPT-OSS 120B (Cerebras)Cerebras$0.35$0.75$0.45$0.54
DeepSeek V4 FlashDeepSeek$0.44$1.32$0.66$0.77
Mistral Large 3Mistral$0.50$1.50$0.75$0.88
Gemini 3.5 Flash LiteGoogle$0.30$2.50$0.85$0.93
Qwen 3.7 PlusQwen$0.80$2.00$1.10$1.30
Qwen 3.8 30BGroq$0.60$3.00$1.20$1.35
Amazon Nova ProAmazon$0.80$3.20$1.40$1.60
Gemini 3.7 FlashGoogle$0.75$3.75$1.50$1.69
Grok 4.3xAI$1.25$2.50$1.56$1.88
DeepSeek V4 ProDeepSeek$1.32$3.96$1.98$2.31
Claude Haiku 4.5Anthropic$1.00$5.00$2.00$2.25
Muse Spark 1.3Meta$1.25$4.25$2.00$2.31
GLM-5.2Z.ai$1.40$4.40$2.15$2.50
GPT-5.6 LunaOpenAI$1.00$6.00$2.25$2.50
GLM 4.7 (Cerebras)Cerebras$2.25$2.75$2.38$2.94
Qwen 3.8 MaxQwen$1.60$6.40$2.80$3.20
Qwen 3.7 MaxQwen$1.60$6.40$2.80$3.20
Grok-4.20 ReasoningxAI$2.00$6.00$3.00$3.50
Grok-4.20xAI$2.00$6.00$3.00$3.50
Grok 4.6xAI$2.00$6.00$3.00$3.50
Grok 4.5xAI$2.00$6.00$3.00$3.50
Gemini 3.6 FlashGoogle$1.50$7.50$3.00$3.38
Mistral Medium 3Mistral$1.50$7.50$3.00$3.38
Claude Sonnet 5Anthropic$2.00$10.00$4.00$4.50
Gemini 3.1 ProGoogle$2.00$12.00$4.50$5.00
GPT-5.6 TerraOpenAI$2.50$15.00$5.63$6.25
Claude Sonnet 4.6Anthropic$3.00$15.00$6.00$6.75
GPT-5.6 SolOpenAI$4.00$20.00$8.00$9.00
Claude Opus 4.8Anthropic$5.00$25.00$10.00$11.25
Claude Fable 5Anthropic$10.00$50.00$20.00$22.50
Claude Opus 5Anthropic$15.00$75.00$30.00$33.75

Three reproducible pricing decisions

Every result is normalized to 100,000 requests. List, prompt-cache, and batch are separate applicable-cost columns; cache uses the provider read multiplier and batch uses a documented 50% scenario only for providers with a reusable rule.

1. Workload-normalized list / cache / batch cost

ModelWorkloadListPrompt-cache readBatch scenario
GPT-5 NanoShort chat$12.00$8.40$8.00
GPT-5 NanoDocument extraction$72.00$36.00$56.00
GPT-5 NanoOutput-heavy drafting$165.00$160.50$85.00
Gemini 2.5 Flash LiteShort chat$16.00$8.80$12.00
Gemini 2.5 Flash LiteDocument extraction$112.00$40.00$96.00
Gemini 2.5 Flash LiteOutput-heavy drafting$170.00$161.00$90.00
GPT-4o MiniShort chat$24.00$13.20$18.00
GPT-4o MiniDocument extraction$168.00$60.00$144.00
GPT-4o MiniOutput-heavy drafting$255.00$241.50$135.00
Grok-3 MiniShort chat$24.00UnavailableUnavailable
Grok-3 MiniDocument extraction$168.00UnavailableUnavailable
Grok-3 MiniOutput-heavy drafting$255.00UnavailableUnavailable
GPT-5.4 NanoShort chat$41.00$26.60$28.50
GPT-5.4 NanoDocument extraction$260.00$116.00$210.00
GPT-5.4 NanoOutput-heavy drafting$520.00$502.00$270.00
Gemini 3.1 Flash LiteShort chat$50.00$32.00$35.00
Gemini 3.1 Flash LiteDocument extraction$320.00$140.00$260.00
Gemini 3.1 Flash LiteOutput-heavy drafting$625.00$602.50$325.00
DeepSeek V4 FlashShort chat$61.60UnavailableUnavailable
DeepSeek V4 FlashDocument extraction$457.60UnavailableUnavailable
DeepSeek V4 FlashOutput-heavy drafting$572.00UnavailableUnavailable
GPT-5 MiniShort chat$60.00$42.00$40.00
GPT-5 MiniDocument extraction$360.00$180.00$280.00
GPT-5 MiniOutput-heavy drafting$825.00$802.50$425.00
Gemini 3.5 Flash LiteShort chat$74.00$52.40$49.00
Gemini 3.5 Flash LiteDocument extraction$440.00$224.00$340.00
Gemini 3.5 Flash LiteOutput-heavy drafting$1030.00$1003.00$530.00
Gemini 2.5 FlashShort chat$74.00$52.40$49.00
Gemini 2.5 FlashDocument extraction$440.00$224.00$340.00
Gemini 2.5 FlashOutput-heavy drafting$1030.00$1003.00$530.00
Gemini 3.7 FlashShort chat$135.00$81.00$97.50
Gemini 3.7 FlashDocument extraction$900.00$360.00$750.00
Gemini 3.7 FlashOutput-heavy drafting$1575.00$1507.50$825.00
Grok 4.3Short chat$150.00UnavailableUnavailable
Grok 4.3Document extraction$1200.00UnavailableUnavailable
Grok 4.3Output-heavy drafting$1125.00UnavailableUnavailable
GPT-5.4 MiniShort chat$150.00$96.00$105.00
GPT-5.4 MiniDocument extraction$960.00$420.00$780.00
GPT-5.4 MiniOutput-heavy drafting$1875.00$1807.50$975.00
Gemini 3.1 FlashShort chat$150.00$96.00$105.00
Gemini 3.1 FlashDocument extraction$960.00$420.00$780.00
Gemini 3.1 FlashOutput-heavy drafting$1875.00$1807.50$975.00
o3-MiniShort chat$176.00$96.80$132.00
o3-MiniDocument extraction$1232.00$440.00$1056.00
o3-MiniOutput-heavy drafting$1870.00$1771.00$990.00
DeepSeek V4 ProShort chat$184.80UnavailableUnavailable
DeepSeek V4 ProDocument extraction$1372.80UnavailableUnavailable
DeepSeek V4 ProOutput-heavy drafting$1716.00UnavailableUnavailable
Claude Haiku 4.5Short chat$180.00$108.00$130.00
Claude Haiku 4.5Document extraction$1200.00$480.00$1000.00
Claude Haiku 4.5Output-heavy drafting$2100.00$2010.00$1100.00
GPT-5.6 LunaShort chat$200.00$128.00$140.00
GPT-5.6 LunaDocument extraction$1280.00$560.00$1040.00
GPT-5.6 LunaOutput-heavy drafting$2500.00$2410.00$1300.00
Grok-3Short chat$240.00UnavailableUnavailable
Grok-3Document extraction$1920.00UnavailableUnavailable
Grok-3Output-heavy drafting$1800.00UnavailableUnavailable
Grok-4.20 ReasoningShort chat$280.00UnavailableUnavailable
Grok-4.20 ReasoningDocument extraction$2080.00UnavailableUnavailable
Grok-4.20 ReasoningOutput-heavy drafting$2600.00UnavailableUnavailable
Grok-4.20Short chat$280.00UnavailableUnavailable
Grok-4.20Document extraction$2080.00UnavailableUnavailable
Grok-4.20Output-heavy drafting$2600.00UnavailableUnavailable
Grok 4.6Short chat$280.00UnavailableUnavailable
Grok 4.6Document extraction$2080.00UnavailableUnavailable
Grok 4.6Output-heavy drafting$2600.00UnavailableUnavailable
Grok 4.5Short chat$280.00UnavailableUnavailable
Grok 4.5Document extraction$2080.00UnavailableUnavailable
Grok 4.5Output-heavy drafting$2600.00UnavailableUnavailable
Gemini 3.6 FlashShort chat$270.00$162.00$195.00
Gemini 3.6 FlashDocument extraction$1800.00$720.00$1500.00
Gemini 3.6 FlashOutput-heavy drafting$3150.00$3015.00$1650.00
Gemini 3.5 FlashShort chat$300.00$192.00$210.00
Gemini 3.5 FlashDocument extraction$1920.00$840.00$1560.00
Gemini 3.5 FlashOutput-heavy drafting$3750.00$3615.00$1950.00
GPT-5Short chat$300.00$210.00$200.00
GPT-5Document extraction$1800.00$900.00$1400.00
GPT-5Output-heavy drafting$4125.00$4012.50$2125.00
GPT-4.1Short chat$320.00$176.00$240.00
GPT-4.1Document extraction$2240.00$800.00$1920.00
GPT-4.1Output-heavy drafting$3400.00$3220.00$1800.00
Claude Sonnet 5Short chat$360.00$216.00$260.00
Claude Sonnet 5Document extraction$2400.00$960.00$2000.00
Claude Sonnet 5Output-heavy drafting$4200.00$4020.00$2200.00
GPT-4oShort chat$400.00$220.00$300.00
GPT-4oDocument extraction$2800.00$1000.00$2400.00
GPT-4oOutput-heavy drafting$4250.00$4025.00$2250.00
Gemini 3.1 ProShort chat$400.00$256.00$280.00
Gemini 3.1 ProDocument extraction$2560.00$1120.00$2080.00
Gemini 3.1 ProOutput-heavy drafting$5000.00$4820.00$2600.00
GPT-5.6 TerraShort chat$500.00$320.00$350.00
GPT-5.6 TerraDocument extraction$3200.00$1400.00$2600.00
GPT-5.6 TerraOutput-heavy drafting$6250.00$6025.00$3250.00
GPT-5.4Short chat$500.00$320.00$350.00
GPT-5.4Document extraction$3200.00$1400.00$2600.00
GPT-5.4Output-heavy drafting$6250.00$6025.00$3250.00
Claude Sonnet 4.6Short chat$540.00$324.00$390.00
Claude Sonnet 4.6Document extraction$3600.00$1440.00$3000.00
Claude Sonnet 4.6Output-heavy drafting$6300.00$6030.00$3300.00
Claude Sonnet 4.5Short chat$540.00$324.00$390.00
Claude Sonnet 4.5Document extraction$3600.00$1440.00$3000.00
Claude Sonnet 4.5Output-heavy drafting$6300.00$6030.00$3300.00
Claude Sonnet 4Short chat$540.00$324.00$390.00
Claude Sonnet 4Document extraction$3600.00$1440.00$3000.00
Claude Sonnet 4Output-heavy drafting$6300.00$6030.00$3300.00
GPT-5.6 SolShort chat$720.00$432.00$520.00
GPT-5.6 SolDocument extraction$4800.00$1920.00$4000.00
GPT-5.6 SolOutput-heavy drafting$8400.00$8040.00$4400.00
Claude Opus 4.8Short chat$900.00$540.00$650.00
Claude Opus 4.8Document extraction$6000.00$2400.00$5000.00
Claude Opus 4.8Output-heavy drafting$10500.00$10050.00$5500.00
Claude Opus 4.7Short chat$900.00$540.00$650.00
Claude Opus 4.7Document extraction$6000.00$2400.00$5000.00
Claude Opus 4.7Output-heavy drafting$10500.00$10050.00$5500.00
Claude Opus 4.6Short chat$900.00$540.00$650.00
Claude Opus 4.6Document extraction$6000.00$2400.00$5000.00
Claude Opus 4.6Output-heavy drafting$10500.00$10050.00$5500.00
Claude Opus 4.5Short chat$900.00$540.00$650.00
Claude Opus 4.5Document extraction$6000.00$2400.00$5000.00
Claude Opus 4.5Output-heavy drafting$10500.00$10050.00$5500.00
GPT-4 TurboShort chat$1400.00$680.00$1100.00
GPT-4 TurboDocument extraction$10400.00$3200.00$9200.00
GPT-4 TurboOutput-heavy drafting$13000.00$12100.00$7000.00
Claude Fable 5Short chat$1800.00$1080.00$1300.00
Claude Fable 5Document extraction$12000.00$4800.00$10000.00
Claude Fable 5Output-heavy drafting$21000.00$20100.00$11000.00
Claude Opus 5Short chat$2700.00$1620.00$1950.00
Claude Opus 5Document extraction$18000.00$7200.00$15000.00
Claude Opus 5Output-heavy drafting$31500.00$30150.00$16500.00
Claude Opus 4.1Short chat$2700.00$1620.00$1950.00
Claude Opus 4.1Document extraction$18000.00$7200.00$15000.00
Claude Opus 4.1Output-heavy drafting$31500.00$30150.00$16500.00
Claude Opus 4Short chat$2700.00$1620.00$1950.00
Claude Opus 4Document extraction$18000.00$7200.00$15000.00
Claude Opus 4Output-heavy drafting$31500.00$30150.00$16500.00
GPT-5.4 ProShort chat$6000.00$3840.00$4200.00
GPT-5.4 ProDocument extraction$38400.00$16800.00$31200.00
GPT-5.4 ProOutput-heavy drafting$75000.00$72300.00$39000.00

2. Workload crossover

WorkloadInput / outputList winnerList cost / 100KApplicability note
Short chat800 / 200GPT-5 Nano$12.00Cache/batch columns below retain unavailable states per model
Document extraction8,000 / 800GPT-5 Nano$72.00Cache/batch columns below retain unavailable states per model
Output-heavy drafting1,000 / 4,000GPT-5 Nano$165.00Cache/batch columns below retain unavailable states per model

Formula: 100,000 × (input tokens × input $/M + output tokens × output $/M) ÷ 1,000,000. A missing cache or batch rule is shown as unavailable, never as free.

3. Dated price-change ledger

ModelVerifiedBefore / after deltaSource
GPT-5 Nano2026-04-06Before value unavailable in current registryFirst-party price record
Gemini 2.5 Flash Lite2026-04-06Before value unavailable in current registryFirst-party price record
GPT-4o Mini2026-04-06Before value unavailable in current registryFirst-party price record
Grok-3 Mini2026-04-06Before value unavailable in current registryFirst-party price record
GPT-5.4 Nano2026-04-06Before value unavailable in current registryFirst-party price record
Gemini 3.1 Flash Lite2026-04-06Before value unavailable in current registryFirst-party price record
DeepSeek V4 Flash2026-08-14Before value unavailable in current registryFirst-party price record
GPT-5 Mini2026-04-06Before value unavailable in current registryFirst-party price record

Historical deltas stay unavailable when the repository has only the current value.

Run an attributed cross-model comparison

Verified 2026-08-14. Formulas use the dated registry; unavailable means no model-specific dated source supports a numeric claim. dated normalized pricing registry.

Batch 14 · rate freshness, unit normalization, and market shock governance

1. Rate-freshness control board

Age bandRowsSource reachabilityLifecycleUnit completenessComparison treatment
0–7 days20Registry source URL presentRegistry lifecycle filter appliedText input/output rows onlyEligible only if units match
8–30 days2Registry source URL presentRegistry lifecycle filter appliedText input/output rows onlyEligible only if units match
31–60 days6Registry source URL presentRegistry lifecycle filter appliedText input/output rows onlyEligible only if units match
61–90 days8Registry source URL presentRegistry lifecycle filter appliedText input/output rows onlyEligible only if units match
90+ days32Registry source URL presentRegistry lifecycle filter appliedText input/output rows onlyExclude all stale rows

Formula: age = verification date − source date. Rows are excluded when stale, unreachable, lifecycle-incompatible, or unit-incomplete; no stale row is silently carried forward.

2. Cross-provider unit-normalization audit

Unit familyRegistry fieldsCompatible comparisonBlended priceVerdict
text input/outputinputPerM + outputPerMLike-for-like token shapeNever blend unlike unitsCompatible
cached tokensUnavailableUnavailableUnavailableUnavailable — no cross-unit blend
image/audio/videoUnavailableUnavailableUnavailableUnavailable — no cross-unit blend
tools/searchUnavailableUnavailableUnavailableUnavailable — no cross-unit blend
storageUnavailableUnavailableUnavailableUnavailable — no cross-unit blend
batchUnavailableUnavailableUnavailableUnavailable — no cross-unit blend
long-context tiersUnavailableUnavailableUnavailableUnavailable — no cross-unit blend

3. Neutral portfolio-shock table

BasketAffected models/providersBase bill0% shock10% shock25% shock50% shockConcentration
text-heavyOpenAI/GPT-5 Nano · Anthropic/Claude Haiku 4.5 · Google/Gemini 2.5 Flash Lite · xAI/Grok-3 Mini$0.02$0.02$0.02$0.03$0.0377.3% top-model share
balancedOpenAI/GPT-5 Nano · Anthropic/Claude Haiku 4.5 · Google/Gemini 2.5 Flash Lite · xAI/Grok-3 Mini$0.03$0.03$0.03$0.04$0.0477.7% top-model share
output-heavyOpenAI/GPT-5 Nano · Anthropic/Claude Haiku 4.5 · Google/Gemini 2.5 Flash Lite · xAI/Grok-3 Mini$0.06$0.06$0.06$0.07$0.0878.0% top-model share

Formula: shocked bill = base compatible bill × (1 + rate shock). This table reports affected rows and spend risk; it does not name a universal cheapest provider.

Verified 2026-08-14. Data owner: Luna. “Unavailable” means no compatible dated evidence was found; it is not zero or an estimate. Source / registry · Run this scenario →

Batch 16 · neutral channel parity, tariff transitions, and missing-unit sensitivity

1. Direct-provider versus cloud-marketplace parity

Model/versionRegionContext tierInput/output/cache/batch/toolCurrency/dateMissing cellsVerdict
Amazon Nova MicroUnavailableUnavailable$0.034999999999999996/M · $0.13999999999999999/M · UnavailableUSD / 2026-06-14UnavailableNo parity verdict
Amazon Nova LiteUnavailableUnavailable$0.060000000000000005/M · $0.24000000000000002/M · UnavailableUSD / 2026-06-14UnavailableNo parity verdict
Muse Spark 1.3 ContributorUnavailableUnavailable$0.1/M · $0.2/M · UnavailableUSD / 2026-08-14UnavailableNo parity verdict
GPT-OSS 20BUnavailableUnavailable$0.075/M · $0.3/M · UnavailableUSD / 2026-04-06UnavailableNo parity verdict
GPT-5 NanoUnavailableUnavailable$0.05/M · $0.4/M · UnavailableUSD / 2026-04-06UnavailableNo parity verdict
Ministral 8BUnavailableUnavailable$0.15/M · $0.15/M · UnavailableUSD / 2026-08-14UnavailableNo parity verdict

Formula / rule: parity requires exact model/version, region, tier, units, currency, and effective date; no cheapest-channel verdict is issued.

2. Tariff effective-date transition board

Model/aliasOld rowNew rowAnnouncementEffectiveObservedBasket deltaStale quarantine
Amazon Nova MicroUnavailableUnavailableUnavailableUnavailable2026-06-14UnavailableQuarantine until joined
Amazon Nova LiteUnavailableUnavailableUnavailableUnavailable2026-06-14UnavailableQuarantine until joined
Muse Spark 1.3 ContributorUnavailableUnavailableUnavailableUnavailable2026-08-14UnavailableQuarantine until joined
GPT-OSS 20BUnavailableUnavailableUnavailableUnavailable2026-04-06UnavailableQuarantine until joined
GPT-5 NanoUnavailableUnavailableUnavailableUnavailable2026-04-06UnavailableQuarantine until joined

Formula / rule: eligible row = observed date ≥ effective date and not superseded; future, stale, and alias-unmatched rows are quarantined before comparison.

3. Missing-unit sensitivity frontier

AttemptKnown billUnknown unitSwitching thresholdNeutral result
textUnavailableUnavailableUnavailableUnavailable
mediaUnavailableUnavailableUnavailableUnavailable
tool/searchUnavailableUnavailableUnavailableUnavailable
storageUnavailableUnavailableUnavailableUnavailable
long-contextUnavailableUnavailableUnavailableUnavailable

Formula / rule: threshold = (competing known bill − known bill) ÷ unknown quantity; if the quantity or compatible unit is absent, result is Unavailable rather than blended.

Verified 2026-08-14. Data owner: Luna. Source / registry: dated repository pricing and provider records. “Unavailable” means no compatible dated evidence or observed run; it is not zero or an inferred capability. Run this evidence scenario →

Batch 18 · billing increments, neutral normalization, and source conflict quarantine

1. Billing-granularity and rounding audit

WorkloadToken rounding / minimumFree allowanceInvoice precisionTheoretical / sourced bill
micro requestUnavailableUnavailableUnavailableUnavailable
short requestUnavailableUnavailableUnavailableUnavailable
long requestUnavailableUnavailableUnavailableUnavailable

Formula / rule: effective cost is Unavailable when billing increment, minimum, allowance, or invoice precision is undocumented; theoretical bills never replace sourced bills.

2. Region-currency-tax comparability board

Model/versionRegion / currencyUnit basis / FX dateTax / surcharge / effectivePretax USD / eligibility
compatible row AUnavailableUnavailableUnavailableUnavailable
compatible row BUnavailableUnavailableUnavailableUnavailable
compatible row CUnavailableUnavailableUnavailableUnavailable

Formula / rule: normalized pretax USD = quoted compatible unit × dated FX, excluding tax; every required field must match before inclusion, and no cheapest provider is named.

3. Price-source conflict resolver

Model/versionProvider docs / consoleMarketplace / registryDisagreement / recencyQuarantine / precedence
row AUnavailableUnavailableUnavailableUnavailable
row BUnavailableUnavailableUnavailableUnavailable
row CUnavailableUnavailableUnavailableUnavailable

Formula / rule: join by model/version and effective date; quarantine disagreement until deterministic precedence (first-party effective documentation, then dated export, then marketplace/registry) is satisfied.

Verified 2026-08-14. Data owner: Luna. Source / registry: dated repository records and matched-run evidence. “Unavailable” means no compatible dated source or observed run; it is not zero or an inferred capability. Run this Batch 18 evidence scenario →

For cross-dataset findings and a citation-ready snapshot, read the latest LLM pricing and performance report.

Cheapest by tier

budget tier
Amazon Nova Micro
$0.061 / M blended tokens
mid tier
GLM-5.1
$1.000 / M blended tokens
frontier tier
GPT-4 Turbo
$15.000 / M blended tokens

Popular head-to-heads

Claude Opus 4 vs Claude Opus 4.8Grok-3 vs Grok 4.3GPT-4o vs GPT-5.6 SolClaude Opus 4.8 vs DeepSeek V4 ProClaude Opus 4.8 vs Gemini 3.1 ProClaude Opus 4.8 vs Gemini 3.7 FlashClaude Opus 4.8 vs GPT-5.6 SolClaude Opus 4.8 vs Grok 4.3Claude Opus 4.8 vs Grok 4.5Claude Opus 4.8 vs Grok 4.6

Methodology & freshness

Prices are hand-verified against each provider's official pricing page and stored with a source URL and verification date. "Blended cost" assumes 3 input tokens per 1 output token. The oldest verified price in this dataset is from 2026-04-06; the newest is from 2026-08-14. Legacy models remain listed (toggle above) because they still see real API traffic — each links to its recommended successor. Cite this: All AI Ask LLM API Pricing Dataset, retrieved 2026-08-14.

FAQ

What does this LLM API pricing comparison include?

This hub compares standard input, cached-input where available, and output prices across 68 models, with verification dates and a consistent blended-cost calculation. For a ranked answer to the separate cheapest-LLM-API question, see the dedicated cost ranking.

Why is output more expensive than input?

Generating a token requires a full forward pass through the model, while processing an input token can be batched and cached. Providers price output 3–6x higher than input to reflect this compute cost.

Do these prices include caching discounts?

No — these are the standard, non-cached list prices published by each provider. Prompt caching (where available) can cut input costs significantly for repeated context.

How often are these prices updated?

We verify prices against each provider's official pricing page and re-check on every model addition. The most recently verified price in this table is dated 2026-08-14; the oldest is 2026-04-06.

What does "blended cost" mean?

Blended cost assumes 3 input tokens for every 1 output token — a rough approximation of typical chat/agent usage — computed as (3 × input price + 1 × output price) / 4.

See real measured cost, not just list price

Run the same prompt across these models side by side and see the actual token cost for your use case.

Try It Free