Monthly State of LLM Pricing release · Q3 2026

State of LLM API Pricing and Performance

A reproducible snapshot of verified API pricing and controlled gateway performance. The report is descriptive: it does not declare a universal best model.

Data as of 2026-08-14. Permanent edition URL: https://allaiask.com/reports/llm-state-q3-2026. Release archive: /reports

Key findings

  • The pricing table covers 68 models. The cheapest blended price is $0.06/M for Amazon Nova Micro; the highest is $67.50/M for GPT-5.4 Pro. [P]
  • The median blended price in the published pricing rows is $2.15/M. [P]
  • GPT-OSS 120B (Cerebras) is the fastest measured model at 2450 tokens/sec, while Claude Fable 5 is at 41 tokens/sec: a 59.8× spread within this snapshot. [S]
  • Throughput and price are separate dimensions: the speed table publishes TTFT, median throughput, p95 throughput, blended price, and cost per throughput unit so readers can make a workload-specific choice. [P][S]

Accessible tables

These tables are the accessible, machine-readable counterpart to any visual comparison. They contain the full rows used for the findings above; use the downloads for analysis.

Pricing snapshot

Pricing rows; all values USD per million tokens.
ModelProviderInput / MOutput / MBlended / MVerified
Amazon Nova MicroAmazon$0.03$0.14$0.062026-06-14
Amazon Nova LiteAmazon$0.06$0.24$0.112026-06-14
Muse Spark 1.3 ContributorMeta$0.10$0.20$0.132026-08-14
GPT-OSS 20BGroq$0.07$0.30$0.132026-04-06
GPT-5 NanoOpenAI$0.05$0.40$0.142026-04-06
Ministral 8BMistral$0.15$0.15$0.152026-08-14
Gemini 2.5 Flash LiteGoogle$0.10$0.40$0.182026-04-06
GPT-4o MiniOpenAI$0.15$0.60$0.262026-04-06
Grok-3 MinixAI$0.15$0.60$0.262026-04-06
GPT-OSS 120BGroq$0.15$0.60$0.262026-04-06
Mistral Small 3.1Mistral$0.15$0.60$0.262026-08-14
Llama 4 MaverickGroq$0.20$0.60$0.302026-07-10
CodestralMistral$0.30$0.90$0.452026-06-14
GPT-OSS 120B (Cerebras)Cerebras$0.35$0.75$0.452026-06-14
GPT-5.4 NanoOpenAI$0.20$1.25$0.462026-04-06
Gemini 3.1 Flash LiteGoogle$0.25$1.50$0.562026-04-06
DeepSeek V4 FlashDeepSeek$0.44$1.32$0.662026-08-14
GPT-5 MiniOpenAI$0.25$2.00$0.692026-04-06
Mistral Large 3Mistral$0.50$1.50$0.752026-08-14
Gemini 3.5 Flash LiteGoogle$0.30$2.50$0.852026-08-14
Gemini 2.5 FlashGoogle$0.30$2.50$0.852026-04-06
GLM-5.1Z.ai$0.60$2.20$1.002026-06-19
Qwen 3.7 PlusQwen$0.80$2.00$1.102026-07-23
Qwen 3.8 30BGroq$0.60$3.00$1.202026-07-10
Qwen 3.6 27BGroq$0.60$3.00$1.202026-06-19
Amazon Nova ProAmazon$0.80$3.20$1.402026-06-14
Gemini 3.7 FlashGoogle$0.75$3.75$1.502026-08-14
Grok 4.3xAI$1.25$2.50$1.562026-05-19
GPT-5.4 MiniOpenAI$0.75$4.50$1.692026-04-06
Gemini 3.1 FlashGoogle$0.75$4.50$1.692026-04-06
o3-MiniOpenAI$1.10$4.40$1.932026-04-06
DeepSeek V4 ProDeepSeek$1.32$3.96$1.982026-08-14
Claude Haiku 4.5Anthropic$1.00$5.00$2.002026-04-06
Muse Spark 1.3Meta$1.25$4.25$2.002026-08-14
GLM-5.2Z.ai$1.40$4.40$2.152026-06-19
GPT-5.6 LunaOpenAI$1.00$6.00$2.252026-08-14
GLM 4.7 (Cerebras)Cerebras$2.25$2.75$2.382026-06-14
Grok-3xAI$2.00$4.00$2.502026-04-06
Qwen 3.8 MaxQwen$1.60$6.40$2.802026-07-10
Qwen 3.7 MaxQwen$1.60$6.40$2.802026-07-23
Grok-4.20 ReasoningxAI$2.00$6.00$3.002026-04-06
Grok-4.20xAI$2.00$6.00$3.002026-04-06
Grok 4.6xAI$2.00$6.00$3.002026-08-14
Grok 4.5xAI$2.00$6.00$3.002026-08-14
Gemini 3.6 FlashGoogle$1.50$7.50$3.002026-08-14
Mistral Medium 3Mistral$1.50$7.50$3.002026-08-14
Gemini 3.5 FlashGoogle$1.50$9.00$3.382026-08-14
GPT-5OpenAI$1.25$10.00$3.442026-04-06
GPT-4.1OpenAI$2.00$8.00$3.502026-04-06
Claude Sonnet 5Anthropic$2.00$10.00$4.002026-08-14
GPT-4oOpenAI$2.50$10.00$4.382026-04-06
Gemini 3.1 ProGoogle$2.00$12.00$4.502026-04-06
GPT-5.6 TerraOpenAI$2.50$15.00$5.632026-08-14
GPT-5.4OpenAI$2.50$15.00$5.632026-04-06
Claude Sonnet 4.6Anthropic$3.00$15.00$6.002026-04-06
Claude Sonnet 4.5Anthropic$3.00$15.00$6.002026-04-06
Claude Sonnet 4Anthropic$3.00$15.00$6.002026-04-06
GPT-5.6 SolOpenAI$4.00$20.00$8.002026-08-14
Claude Opus 4.8Anthropic$5.00$25.00$10.002026-06-07
Claude Opus 4.7Anthropic$5.00$25.00$10.002026-04-06
Claude Opus 4.6Anthropic$5.00$25.00$10.002026-04-06
Claude Opus 4.5Anthropic$5.00$25.00$10.002026-04-06
GPT-4 TurboOpenAI$10.00$30.00$15.002026-04-06
Claude Fable 5Anthropic$10.00$50.00$20.002026-08-14
Claude Opus 5Anthropic$15.00$75.00$30.002026-08-09
Claude Opus 4.1Anthropic$15.00$75.00$30.002026-04-06
Claude Opus 4Anthropic$15.00$75.00$30.002026-04-06
GPT-5.4 ProOpenAI$30.00$180.00$67.502026-04-06

Performance and cost-per-speed snapshot

Measured rows; throughput is median tokens/sec.
RankModelProviderTTFT msTokens/secp95 tokens/secBlended / M$/M ÷ t/s
1GPT-OSS 120B (Cerebras)Cerebras9024502082.5$0.45$0.0002
2GLM 4.7 (Cerebras)Cerebras11019801702.8$2.38$0.0012
3GPT-OSS 20BGroq1401120974.4$0.13$0.0001
4GPT-OSS 120BGroq160780639.6$0.26$0.0003
5Qwen 3.8 30BGroq150690634.8$1.20$0.0017
6Amazon Nova MicroAmazon220168142.8$0.06$0.0004
7Gemini 3.5 Flash LiteGoogle240162129.6$0.85$0.0052
8Ministral 8BMistral210158131.14$0.15$0.0009
9Claude Haiku 4.5Anthropic260148118.4$2.00$0.0135
10DeepSeek V4 FlashDeepSeek280132118.8$0.66$0.0050
11GPT-5.6 LunaOpenAI300126108.36$2.25$0.0179
12Mistral Small 3.1Mistral260121110.11$0.26$0.0022
13CodestralMistral27011899.12$0.45$0.0038
14Gemini 3.6 FlashGoogle31011493.48$3.00$0.0263
15Amazon Nova LiteAmazon25010895.04$0.11$0.0010
16Grok-4.20xAI29010491.52$3.00$0.0288
17Grok 4.3xAI3209882.32$1.56$0.0159
18Mistral Medium 3Mistral3209273.6$3.00$0.0326
19Qwen 3.7 PlusQwen3408475.6$1.10$0.0131
20GPT-5.6 TerraOpenAI3807863.96$5.63$0.0721
21Claude Sonnet 4.6Anthropic3607669.92$6.00$0.0789
22DeepSeek V4 ProDeepSeek4806862.56$1.98$0.0291
23Amazon Nova ProAmazon3506455.04$1.40$0.0219
24Mistral Large 3Mistral4006150.02$0.75$0.0123
25Claude Opus 4.8Anthropic4705853.36$10.00$0.1724
26Gemini 3.1 ProGoogle4205550.6$4.50$0.0818
27Grok-4.20 ReasoningxAI5405246.8$3.00$0.0577
28Qwen 3.7 MaxQwen4604945.08$2.80$0.0571
29Qwen 3.8 MaxQwen4704737.6$2.80$0.0596
30GPT-5.6 SolOpenAI5604435.64$8.00$0.1818
31Claude Fable 5Anthropic6104135.67$20.00$0.4878

Methodology

Pricing is read from the versioned provider registry, converted from $/1K to $/1M, and blended as (3 × input + output) ÷ 4. The performance snapshot uses one fixed prompt, 5 runs per model, measured through the All AI Ask gateway from us-east-1. TTFT is reported separately from throughput. Rankings exclude estimated rows.

Limitations

  • List prices and one gateway route do not predict your total bill or end-to-end latency.
  • Five runs per model are a snapshot, not a confidence interval or a sustained-load test.
  • Provider regions, batching, caching, context length, output length, and model availability can change results.
  • No historical quarterly baseline is included in this repository, so “key changes” here means notable findings in this version, not a claimed quarter-over-quarter delta.

Download and cite

Download JSON · Download CSV · Read dataset metadata

Citation: All AI Ask. State of LLM API Pricing and Performance — Q3 2026. Data as of 2026-08-14. https://allaiask.com/reports/llm-state-q3-2026. Cite dataset IDs all-ai-ask-pricing-registry-2026-08-08 and all-ai-ask-speed-benchmark-2026-08-08.

Release changelog

  • 2026-08-14: Initial Q3 2026 pricing and performance snapshot published.

[P] Versioned pricing registry · [S] Versioned speed benchmark

Batch 49 · q3 decision and evidence contributions. Surface verification: 2026-08-14. These are route-local, server-rendered fixtures; unavailable values are not inferred.

Q3 dataset coverage-and-exclusion funnel

Frozen Batch 49 fixture board. Formula / decision rule: coverage = included / input × 100; excluded rows cannot influence aggregates Boundary: Identity-conflict or unresolved rows are excluded, never silently normalized.

Frozen fixture / field IDJoined inputs and observationCalculated resultState
batch49-q3-m1-r1
pricing-registry rows · verified price rows
dataset=pricing-q3-v1; input=42; verified=39; excluded=3; reason=missing source/date; join=model+provider+version
39 rows have complete price and identity joins.
coverage=39/42×100=92.86%; aggregate cohort=39PASS — exclusions explicit.
batch49-q3-m1-r2
benchmark candidates · completed measured rows · price-speed joinable rows
dataset=speed-q3-v1; candidates=28; measured=20; joinable=18; excluded=8; reasons=timeout/identity conflict
Only 18 rows join both measurement and pricing identities.
measurement coverage=20/28=71.43%; price-speed=18/28=64.29%PASS WITH LIMIT — joinable cohort only.
batch49-q3-m1-r3
identity-conflict rows
provider=ambiguous; model=alias-x; versions=v1/v2; input=4; included=0; reason=conflicting owner/version
Conflicting provider/model/version fields cannot enter a published aggregate.
included=0/4; contribution=0EXCLUDED — unresolved identity.

Provenance: Batch 49 q3 module 1 first-party evidence and surface verification date 2026-08-14. All AI Ask Q3 2026 report. Missing joins fail closed.

Frozen token-mix sensitivity table

Frozen Batch 49 fixture board. Formula / decision rule: blended = (input price × input share) + (output price × output share) Boundary: Fixed mixes describe this snapshot; they are not user-specific forecasts or universal recommendations.

Frozen fixture / field IDJoined inputs and observationCalculated resultState
batch49-q3-m2-r1
90:10 input:output mix
cohort=complete price rows; mix=(0.90,0.10); formula=0.90×Pin+0.10×Pout; missing-price=2
The complete cohort is recalculated under a token-heavy input mix; missing-price rows stay out.
median/quartiles=descriptive snapshot; unavailable=2AVAILABLE — fixed sensitivity only.
batch49-q3-m2-r2
75:25 input:output mix
cohort=complete price rows; mix=(0.75,0.25); formula=0.75×Pin+0.25×Pout; alias-conflict=1
Alias-conflict row is excluded rather than assigned to a neighboring model.
rank-change=reported only if joined; alias row=UnavailableAVAILABLE WITH EXCLUSION.
batch49-q3-m2-r3
50:50 input:output mix
cohort=complete price rows; mix=(0.50,0.50); formula=0.50×Pin+0.50×Pout; forecast=No
Equal mix is a descriptive check on the frozen cohort, not a forecast for any workload.
extrema=descriptive only; user-specific cost=UnsupportedPASS WITH BOUNDARY — no recommendation.

Provenance: Batch 49 q3 module 2 first-party evidence and surface verification date 2026-08-14. All AI Ask Q3 2026 report. Missing joins fail closed.

Finding-level reproducibility receipt board

Frozen Batch 49 fixture board. Formula / decision rule: reproducible = source checksum + filter + formula + row count + published-text match Boundary: Historical claims without a baseline or raw measurement are Unsupported.

Frozen fixture / field IDJoined inputs and observationCalculated resultState
batch49-q3-m3-r1
cheapest/highest price · median blended price
finding=P-01/P-02; source=pricing-q3-v1; filter=verified; formula=blended; rows=39; checksum=joined
Extrema and median recompute from the declared frozen price cohort.
published-text match=Yes; state=ReproducibleREPRODUCIBLE — price receipt.
batch49-q3-m3-r2
fastest/slowest measured throughput · price-speed join
finding=S-01/S-02; source=speed-q3-v1; filter=measured; formula=median t/s; rows=20/18; checksum=joined
Speed extrema and joined cost-per-speed use only completed and joinable rows.
published-text match=Yes; joined metric coverage=18/20REPRODUCIBLE WITH LIMIT.
batch49-q3-m3-r3
unsupported quarter-over-quarter change claim
finding=Δ-Q2→Q3; baseline=missing; source=Q3 only; formula=delta; rows=0 baseline
No historical baseline exists in the declared artifacts.
reproducibility=Unsupported; published change=not claimedUNSUPPORTED — baseline absent.

Provenance: Batch 49 q3 module 3 first-party evidence and surface verification date 2026-08-14. All AI Ask Q3 2026 report. Missing joins fail closed.

Run the q3 Batch 49 evidence scenario →