LLM API Providers Compared — 12 Providers, 39 Models, Updated August 2026

Raw dataset: data.json. Cite this: All AI Ask LLM API Provider Capability Dataset, retrieved 2026-08-14.

Amazon is the cheapest provider by entry-level current-model price. Cerebras is the fastest by median measured throughput, at 2215 tokens/sec. Google has the widest context window, at 2M tokens.

How to read this: "lab" providers train the weights they serve; "host" providers (Groq, Cerebras) serve someone else's open-weight models on their own inference hardware — their differentiator is speed, not model quality.

ProviderKindModelsPrice range /MMax contextMedian tok/sOpenAI-compatiblePrompt cachingBatch discountFree tier / limitsData residencyNext retirement
OpenAIlab3 (+12 legacy)$2.25–$8.001M78YesYes50%
No free tier
No free API tier published; API usage is billed under the account's usage tier.
Official terms
US by default; EU data residency available on enterprise agreements
Anthropiclab6 (+7 legacy)$2.00–$30.001M67PartialYes50%
No free tier
No free API tier published; API usage requires an enabled billing account.
Official terms
Not documented
Googlelab4 (+5 legacy)$0.85–$4.502M114PartialYes50%
Free tier with daily request cap on Google AI Studio
Free-tier requests and tokens vary by model and project; Google publishes the current quota table.
Official terms
Global by default; Vertex AI offers selectable regional endpoints
xAIlab5 (+2 legacy)$1.56–$3.001M98YesNot documentedNot documented
Free starting credits for new accounts
Promotional credits are limited to eligible new accounts; amount and expiry vary by account.
Official terms
Not documented
DeepSeeklab2$0.66–$1.981M100YesNot documentedNot documented
No free tier
No free API tier published; API usage is billed at the listed token rates.
Official terms
Not documented
Mistrallab5$0.15–$3.00256K118YesNot documented50%
Free tier with rate-limited experimentation
Experiment access is rate-limited; current limits depend on account and model.
Official terms
EU-hosted by default
Groqhost3 (+2 legacy)$0.13–$1.20131K780YesNot documented50%
Free tier with per-minute and per-day token caps
Per-minute and per-day request/token caps vary by model and account tier.
Official terms
US
Cerebrashost2$0.45–$2.38200K2215YesNot documentedNot documented
Free tier with a daily token cap
Free access is subject to a daily token cap and account quota.
Official terms
US
Qwenlab3$1.10–$2.80256K49PartialNot documentedNot documented
Free quota for new Alibaba Cloud accounts
Quota is limited to eligible new accounts and varies by model, region, and account.
Official terms
Singapore/international region via DashScope Intl; mainland China served from a separate region
Amazonlab3$0.06–$1.40300K108NoNot documented50%
No free tier
No provider-wide Nova API free tier is published; AWS account promotions, if any, are separate.
Official terms
Selectable AWS region
Z.ailab1 (+1 legacy)$2.15–$2.151MYesNot documentedNot documented
Free trial credits for new accounts
Trial credits are limited to eligible new accounts; amount and expiry vary by account.
Official terms
Not documented
Metalab2$0.13–$2.001.0MYesNot documentedNot documented
No free tier
No published free tier for the Model API; the Contributor tier is a steep discount, not a free tier.
Official terms
Not documented

Every cell is a sourced fact or an explicit "Not documented" — never an inferred value. Full sourcing and verification dates on each provider's hub.

Provider hubs

OpenAI
3 current models · lab
Anthropic
6 current models · lab
Google
4 current models · lab
xAI
5 current models · lab
DeepSeek
2 current models · lab
Mistral
5 current models · lab
Groq
3 current models · host
Cerebras
2 current models · host
Qwen
3 current models · lab
Amazon
3 current models · lab
Z.ai
1 current model · lab
Meta
2 current models · lab

One key, every provider

You don't need each provider's SDK, key, billing relationship, or rate-limit tier — the same call reaches all 12 of them.

Try It Free

FAQ

Which LLM provider is cheapest?

By entry-level current-model price, Amazon is the cheapest provider on this site, with Amazon Nova Micro at $0.06/M blended tokens.

Which providers are OpenAI-compatible?

OpenAI, xAI, DeepSeek, Mistral, Groq, Cerebras, Z.ai, Meta expose a fully OpenAI-compatible endpoint. Anthropic, Google, Qwen offer a partial/beta compatibility layer. Amazon do not.

Do I need a separate API key for each provider?

Yes, if you call each provider directly — each has its own key, billing relationship, and rate-limit tier. Routing every model through All AI Ask removes that: one key reaches all of them.

Which providers offer a free tier?

Google, xAI, Mistral, Groq, Cerebras, Qwen, Z.ai publish a free tier. The rest explicitly publish no free tier as of their verification date.

Which provider has the largest context window?

Google has the widest context window among current models, at 2M tokens.

Provider custody taxonomy, hard-gate shortlist, and concentration planner

Batch 47 decision and evidence surface · verified 2026-08-14 · exact route allowlist match: /llm-providers.

Provider-class and custody normalizer

Frozen Batch 47 provider-directory fixture — Sitemap providers classified by model owner, API seller, serving host, and support party; first-party evidence checked 2026-08-14; unresolved joins render Unavailable.

Method: every calculation uses only the joined fields shown in its row; missing joins fail closed. First-party source: OpenAI API documentation. Verified: 2026-08-14.

Fixture / field IDNamed inputsObservationFormula / boundaryDecision
lab-trained / direct API
batch47-provider-directory-m1-r1
model owner=OpenAI; API seller=OpenAI; serving host=OpenAI; endpoint/region=joined; artifact=Unavailable; billing/support=OpenAIOwner, seller, and host align; artifact visibility remains Unavailable for the managed surface.class = lab-trained/direct when owner = seller = serving host
Boundary: Managed artifact opacity is not a model defect.
PASS — exact layer retained.
lab-trained / cloud-served
batch47-provider-directory-m1-r2
model owner=Meta; API seller=cloud platform; serving host=cloud platform; region=eu-west-1; revision=joined; billing=sellerModel owner differs from API seller and host; support follows the seller.class = lab-trained/cloud-served when model owner ≠ API seller = host
Boundary: Provider counts must not collapse lab and host identities.
PASS — split custody displayed.
third-party host + private artifact
batch47-provider-directory-m1-r3
model owner=Mistral/unknown; API seller=inference host/private team; host=host/private; hash=joined/private; precision=Unavailable; support=splitThird-party host and private artifact have different support and evidence boundaries.class = exact declared layer; missing precision remains Unavailable
Boundary: A private artifact cannot inherit the host provider promise.
UNAVAILABLE — precision/support edge unresolved.

Hard-requirement shortlist compiler

Frozen Batch 47 provider-directory fixture — Frozen workload gates with pass/fail/unknown outcomes before ranking; first-party evidence checked 2026-08-14; unresolved joins render Unavailable.

Method: every calculation uses only the joined fields shown in its row; missing joins fail closed. First-party source: OpenAI API documentation. Verified: 2026-08-14.

Fixture / field IDNamed inputsObservationFormula / boundaryDecision
text extraction + multimodal RAG
batch47-provider-directory-m2-r1
gates=schema/context/citations/media; provider evidence=3/4; eligible set=Unavailable; disqualifier=media custodyText extraction passes; multimodal RAG has unknown media custody, so the set is not eligible.eligible = every hard gate passes; unknown ≠ pass
Boundary: Price and speed are not ranked before every gate passes.
UNAVAILABLE — media custody gate.
five-tool agent + EU-controlled workload
batch47-provider-directory-m2-r2
gates=5 tools/EU region/logging; pass=2; fail=1; unknown=1; eligible=none; disqualifier=tool effect keyTool effect-key failure disqualifies the candidate; EU logging remains unknown.shortlist = providers where fail=0 ∧ unknown=0 across hard gates
Boundary: A model-quality pass cannot waive an effect-control failure.
FAIL — no eligible provider.
500K synthesis + high-burst open-weight + AWS-native app
batch47-provider-directory-m2-r3
gates=context/burst/artifact/IAM; context=unknown; burst=pass; artifact=pass; IAM=pass; eligible=UnavailableThree gates pass, but 500K context evidence is unknown; no provider is promoted.gate coverage = pass / declared = 3/4; eligibility requires 4/4
Boundary: Partial gate coverage cannot produce a winner.
UNAVAILABLE — context gate unresolved.

Portfolio concentration and fallback planner

Frozen Batch 47 provider-directory fixture — Workload shares, owner/host concentration, correlated controls, and fallback qualification; first-party evidence checked 2026-08-14; unresolved joins render Unavailable.

Method: every calculation uses only the joined fields shown in its row; missing joins fail closed. First-party source: OpenAI API documentation. Verified: 2026-08-14.

Fixture / field IDNamed inputsObservationFormula / boundaryDecision
single-provider architecture · six workload roles
batch47-provider-directory-m3-r1
shares=text 40%; RAG 20%; code 15%; tools 10%; media 10%; batch 5%; owner concentration=100%; host=100%All roles share one model owner and serving provider; fallback is not qualified.concentration = Σ(share²); owner=1.00; serving provider=1.00
Boundary: A single-provider plan is a concentration measure, not a quality verdict.
FAIL — fallback absent.
lab-plus-host + two-lab architecture
batch47-provider-directory-m3-r2
owner shares=60/40; serving shares=50/30/20; region=US/EU; common control plane=partial; fallback=2/6 rolesOwner and serving-provider concentration differ; only two roles have replay-qualified fallbacks.uncovered roles = declared roles − qualified fallback roles = 6 − 2 = 4
Boundary: Provider diversity does not guarantee workload fallback.
UNAVAILABLE — four roles uncovered.
region-diverse + private-fallback architecture
batch47-provider-directory-m3-r3
roles=6; regions=US/EU/private; owner shares=45/35/20; serving shares=40/30/20/10; correlated control=none joined; fallback=6/6All six roles have a declared fallback; correlated-control evidence is still only partially joined.fallback coverage = 6/6 = 100%; concentration reported separately by owner and host
Boundary: Coverage does not prove independence when common controls are unknown.
PASS WITH UNAVAILABLE EDGE — independence review remains.

Route owner CTA: test this decision path at All AI Ask. This surface does not transfer evidence to another host, model, region, tier, artifact, or scenario.