← All providers

Groq API Pricing, Models & Rate Limits (2026)

Groq doesn't train models — it serves open-weight models (OpenAI's gpt-oss and Alibaba's Qwen among them) on its own LPU inference hardware. The pitch is raw throughput: Groq is consistently among the fastest tokens-per-second on this site's speed benchmarks for the weights it hosts.

How much does the Groq API cost?

Groq API pricing is usage-based and varies by model and token volume. Use the current model table and provider documentation to estimate a workload before committing to production.

Verified 2026-08-14 source

For cross-provider quota units and fixed-workload capacity, see the LLM API rate-limit comparison; this page remains the authoritative owner for Groq provider facts.

Groq product vs API

Groq consumer access and API billing are separate surfaces; check the provider documentation for current account terms.

Three decisions unique to Groq

Groq current-model price mechanics

Current modelInputCached inputOutputBatchVerified
GPT-OSS 20B$0.075/MUnavailable — no model cache rate$0.300/M50% off eligible Batch API2026-04-06
GPT-OSS 120B$0.150/MUnavailable — no model cache rate$0.600/M50% off eligible Batch API2026-04-06
Llama 4 Maverick$0.200/MUnavailable — no model cache rate$0.600/M50% off eligible Batch API2026-07-10
Qwen 3.8 30B$0.600/MUnavailable — no model cache rate$3.000/M50% off eligible Batch API2026-07-10
Qwen 3.6 27B$0.600/MUnavailable — no model cache rate$3.000/M50% off eligible Batch API2026-06-19

Messages API vs OpenAI compatibility map

ChoiceDecision ruleEvidence
BillingChatGPT plan never includes API creditsSeparate metered API account
Input/cached/outputToken prices are model rowsUse calculator for workload totals
Limits/authPer-model requests/minute and tokens/minute caps by tier · Bearer API keyVerify before production

Adoption map: what is documented versus unavailable

DimensionRecorded valueDecision consequence
Try Groq side by side →

Verified 2026-08-14. dated provider pricing/source

Current models
3
Legacy models
2
Price range /M
$0.13–$1.20
Max context
131K
Median tok/s
780
Next retirement

Groq model pricing

Compare Groq models by input, output, and blended token cost below.

ModelInput /MOutput /MBlended /M
GPT-OSS 20B$0.07$0.30$0.13
GPT-OSS 120B$0.15$0.60$0.26
Qwen 3.8 30B$0.60$3.00$1.20

Pricing values are registry-backed and were most recently verified on 2026-07-10. Sources: https://console.groq.com/docs/models, https://groq.com/pricing. Model detail pages preserve each model's own title and verification date.

* Blended comparison assumes 3 input tokens for every output token; it is not the provider's billing unit.

2 legacy Groq models
Llama 4 Maverick$0.30/M blended
Qwen 3.6 27B$1.20/M blended

Speed

Fastest measured Groq model is GPT-OSS 20B at 1120 tokens/sec (140ms TTFT), median across measured Groq models is 780 tokens/sec. See the full speed benchmark methodology.

Best for

Muse Spark 1.3 Contributor is our pick for CodingMuse Spark 1.3 Contributor is our pick for Structured Data ExtractionMuse Spark 1.3 Contributor is our pick for Writing & ContentGLM-5.2 is our pick for Math & ReasoningMuse Spark 1.3 Contributor is our pick for Agents & Tool Use
What it will cost →
Groq's 5 priced models, ranked by verbosity-adjusted monthly cost, not list rate.

Related Groq pages

Groq alternatives →All LLM API pricing →Groq speed benchmarks →Groq cost calculator →

Build with Groq

Get a Groq API key →Groq rate limits →

Groq implementation details

Verified 2026-08-14 against source.

Groq publishes its current authentication, limits, and data-handling details in the linked documentation.

OpenAI-compatibleYes
API base URLhttps://api.groq.com/openai/v1
Auth modelBearer API key
Prompt cachingNot documented
Batch discount50%
Free tierFree tier with per-minute and per-day token caps
Free-tier limitsPer-minute and per-day request/token caps vary by model and account tier.
Free-tier expiryNot published
Rate-limit modelPer-model requests/minute and tokens/minute caps by tier
Data residencyUS
Trains on API dataNot documented
SLA publishedNo
DocsOfficial pricingStatus pageFree-tier terms

Lifecycle

Groq has 2 legacy models still routable and 4 retired models. Full dates and successors on the model deprecation tracker.

Switching to and from Groq

The closest parity-aware alternative to Qwen 3.8 30B ($1.20/M) outside Groq is GPT-OSS 120B (Cerebras) ($0.45/M, -62.5%) — a config migration. Biggest gap: no vision input.
The closest parity-aware alternative to GPT-OSS 120B ($0.26/M) outside Groq is GPT-OSS 120B (Cerebras) ($0.45/M, +71.4%) — a config migration. Biggest gap: loses the 50% batch discount.
Full Groq alternatives comparison →

Calling Groq through All AI Ask

curl https://allaiask.com/api/v1/chat \
  -H "Authorization: Bearer $ALLAIASK_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "gpt-oss-20b", "messages": [{"role": "user", "content": "Hello"}]}'

FAQ

Is Groq OpenAI-compatible?

Yes — Groq's API base (https://api.groq.com/openai/v1) accepts the OpenAI SDK request/response shape, so existing OpenAI client code works with a base-URL and key swap.

Does Groq support prompt caching?

Not documented as of 2026-08-14 — we did not find a published prompt-caching feature for Groq. If that changes, this page updates.

Does Groq have a free tier?

Yes — Free tier with per-minute and per-day token caps. Per-minute and per-day request/token caps vary by model and account tier.

How much does the Groq API cost?

Current Groq models range from $0.13 to $1.20 per million blended tokens (3:1 input:output). Full per-model pricing is in the table below.

Where is Groq API data hosted?

US

Groq hosted identity, bucket admission, and measured service quality

Batch 47 decision and evidence surface · verified 2026-08-14 · exact route allowlist match: /llm-providers/groq.

Hosted-model provenance and alias ledger

Frozen Batch 47 groq fixture — Groq hosted roster, aliases, revisions, owner withdrawal, and artifact visibility; first-party evidence checked 2026-08-14; unresolved joins render Unavailable.

Method: every calculation uses only the joined fields shown in its row; missing joins fail closed. First-party source: Groq API documentation. Verified: 2026-08-14.

Fixture / field IDNamed inputsObservationFormula / boundaryDecision
current hosted roster row
batch47-groq-m1-r1
owner=Meta; Groq model ID=llama-3.3-70b-versatile; alias=joined; revision=Unavailable; artifact=Unavailable; lifecycle=currentOwner and endpoint ID are visible; revision and artifact are not exposed.cross-date join = owner ∧ Groq ID ∧ revision ∧ lifecycle window
Boundary: Current row cannot absorb an older benchmark with unresolved revision.
UNAVAILABLE — cross-date identity blocked.
frozen alias change + revision replacement
batch47-groq-m1-r2
old alias=llama-3.1; new alias=llama-3.3; change date=2026-07-19; dependent workloads=12; migration owner=platformAlias change is recorded; old and new revisions are not treated as one model.join allowed only when alias transition ∧ exact revision evidence ∧ date window
Boundary: Alias continuity is not revision continuity.
PASS WITH MIGRATION — dependent workloads flagged.
owner withdrawal + hosted-artifact ambiguity
batch47-groq-m1-r3
owner=withdrawn; endpoint ID=legacy; artifact=Unavailable; first evidence=2026-05-03; last evidence=2026-08-14; lifecycle=conflictThe legacy endpoint has conflicting lifecycle signals and no artifact join.lifecycle result = conflict when authoritative signals disagree
Boundary: Do not reuse dependent results after an unresolved withdrawal.
UNAVAILABLE — lifecycle conflict.

Multi-bucket admission calculator

Frozen Batch 47 groq fixture — RPM, TPM, RPD, TPD, tier, reserve, and fail-closed admission fixtures; first-party evidence checked 2026-08-14; unresolved joins render Unavailable.

Method: every calculation uses only the joined fields shown in its row; missing joins fail closed. First-party source: Groq API documentation. Verified: 2026-08-14.

Fixture / field IDNamed inputsObservationFormula / boundaryDecision
60 short requests
batch47-groq-m2-r1
tier=dev; RPM=30; TPM=120K; demand=60 req/9K tok; RPD=Unavailable; retry-after=headerRPM binds first: 30 admitted and 30 deferred; daily capacity is not inferred.admitted = min(requests, RPM, floor(TPM / tokens per request)) = min(60,30,13)=13 per minute
Boundary: Missing RPD does not become unlimited daily capacity.
PASS WITH DEFERRED — RPM/TPM bound.
ten 32K prompts + four long outputs
batch47-groq-m2-r2
model=tier-joined; prompt reserve=320K; output reserve=32K; TPM=Unavailable; TPD=Unavailable; binding bucket=unknownPrompt and output reserve are declared, but token buckets are not documented for this account/model pair.admission = Unavailable when binding bucket is unknown; no arithmetic substitution
Boundary: One model limit cannot be generalized across the provider.
UNAVAILABLE — token bucket missing.
20 parallel tools + one batch-shaped workload
batch47-groq-m2-r3
parallel=20; tool tokens=joined; batch bucket=Unavailable; account tier=prod; retry authority=provider docsTool requests join; batch bucket and deferred-count rule are not documented.admitted count requires every applicable bucket; unknown bucket closes result
Boundary: A documented retry header does not establish batch admission.
UNAVAILABLE — batch bucket unknown.

Groq service-quality measurement card

Frozen Batch 47 groq fixture — Matched endpoint/model/tier observations separated from provider claims; first-party evidence checked 2026-08-14; unresolved joins render Unavailable.

Method: every calculation uses only the joined fields shown in its row; missing joins fail closed. First-party source: Groq API documentation. Verified: 2026-08-14.

Fixture / field IDNamed inputsObservationFormula / boundaryDecision
100-token answer + 4K output
batch47-groq-m3-r1
endpoint/model/tier=joined; n=20; queue=Unavailable; TTFT=176ms median; generation=412 tok/s; total=2.1s; accepted=20/20Acceptance coverage is complete; queue time is unavailable, so total decomposition is bounded.accepted coverage = 20/20 = 100%; total decomposition remains Unavailable without queue
Boundary: Measured throughput is not a model-quality conclusion.
PASS WITH UNAVAILABLE QUEUE — measured card.
strict schema + three-tool loop
batch47-groq-m3-r2
schema=sha256:91c2; tool IDs=3/3; request IDs=joined; retries=1; errors=1; observation window=2026-08-14One tool error retries and final schema passes; retry is counted, not hidden.accepted completion = final checker pass after joined retry; retry rate = 1/20 = 5%
Boundary: A passing final answer does not erase tool failure.
PASS WITH RETRY — scoped observation.
long-context + burst
batch47-groq-m3-r3
context=64K; concurrent=64; accepted=51; errors=13; TTFT=Unavailable; claim source=Groq docs; measured window=joinedBurst acceptance is 51/64; provider speed claim remains separate from the measured error rate.completion coverage = 51/64 = 79.6875%; no rounding to a pass
Boundary: Advertised speed cannot cover burst failures.
FAIL — burst acceptance below full coverage.

Route owner CTA: test this decision path at All AI Ask. This surface does not transfer evidence to another host, model, region, tier, artifact, or scenario.

Try Groq for free

Run real prompts against every current Groq model, and every other provider on this site, in one workspace.

Try It Free