Groq API Pricing, Models & Rate Limits (2026)
Groq doesn't train models — it serves open-weight models (OpenAI's gpt-oss and Alibaba's Qwen among them) on its own LPU inference hardware. The pitch is raw throughput: Groq is consistently among the fastest tokens-per-second on this site's speed benchmarks for the weights it hosts.
How much does the Groq API cost?
Groq API pricing is usage-based and varies by model and token volume. Use the current model table and provider documentation to estimate a workload before committing to production.
For cross-provider quota units and fixed-workload capacity, see the LLM API rate-limit comparison; this page remains the authoritative owner for Groq provider facts.
Groq product vs API
Groq consumer access and API billing are separate surfaces; check the provider documentation for current account terms.
Three decisions unique to Groq
Groq current-model price mechanics
| Current model | Input | Cached input | Output | Batch | Verified |
|---|
| GPT-OSS 20B | $0.075/M | Unavailable — no model cache rate | $0.300/M | 50% off eligible Batch API | 2026-04-06 |
|---|
| GPT-OSS 120B | $0.150/M | Unavailable — no model cache rate | $0.600/M | 50% off eligible Batch API | 2026-04-06 |
|---|
| Llama 4 Maverick | $0.200/M | Unavailable — no model cache rate | $0.600/M | 50% off eligible Batch API | 2026-07-10 |
|---|
| Qwen 3.8 30B | $0.600/M | Unavailable — no model cache rate | $3.000/M | 50% off eligible Batch API | 2026-07-10 |
|---|
| Qwen 3.6 27B | $0.600/M | Unavailable — no model cache rate | $3.000/M | 50% off eligible Batch API | 2026-06-19 |
|---|
Messages API vs OpenAI compatibility map
| Choice | Decision rule | Evidence |
|---|
| Billing | ChatGPT plan never includes API credits | Separate metered API account |
|---|
| Input/cached/output | Token prices are model rows | Use calculator for workload totals |
|---|
| Limits/auth | Per-model requests/minute and tokens/minute caps by tier · Bearer API key | Verify before production |
|---|
Adoption map: what is documented versus unavailable
| Dimension | Recorded value | Decision consequence |
|---|
Try Groq side by side →Verified 2026-08-14. dated provider pricing/source →
Price range /M
$0.13–$1.20
Groq model pricing
Compare Groq models by input, output, and blended token cost below.
Pricing values are registry-backed and were most recently verified on 2026-07-10. Sources: https://console.groq.com/docs/models, https://groq.com/pricing. Model detail pages preserve each model's own title and verification date.
* Blended comparison assumes 3 input tokens for every output token; it is not the provider's billing unit.
2 legacy Groq models
Speed
Best for
Build with Groq
Groq implementation details
Verified 2026-08-14 against source.
Groq publishes its current authentication, limits, and data-handling details in the linked documentation.
| OpenAI-compatible | Yes |
| API base URL | https://api.groq.com/openai/v1 |
| Auth model | Bearer API key |
| Prompt caching | Not documented |
| Batch discount | 50% |
| Free tier | Free tier with per-minute and per-day token caps |
| Free-tier limits | Per-minute and per-day request/token caps vary by model and account tier. |
| Free-tier expiry | Not published |
| Rate-limit model | Per-model requests/minute and tokens/minute caps by tier |
| Data residency | US |
| Trains on API data | Not documented |
| SLA published | No |
Lifecycle
Groq has 2 legacy models still routable and 4 retired models. Full dates and successors on the model deprecation tracker.
Switching to and from Groq
The closest parity-aware alternative to
GPT-OSS 120B ($0.26/M) outside Groq is
GPT-OSS 120B (Cerebras) ($0.45/M, +71.4%) — a
config migration. Biggest gap: loses the 50% batch discount.
Full Groq alternatives comparison →Calling Groq through All AI Ask
curl https://allaiask.com/api/v1/chat \
-H "Authorization: Bearer $ALLAIASK_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "gpt-oss-20b", "messages": [{"role": "user", "content": "Hello"}]}'FAQ
Is Groq OpenAI-compatible?
Yes — Groq's API base (https://api.groq.com/openai/v1) accepts the OpenAI SDK request/response shape, so existing OpenAI client code works with a base-URL and key swap.
Does Groq support prompt caching?
Not documented as of 2026-08-14 — we did not find a published prompt-caching feature for Groq. If that changes, this page updates.
Does Groq have a free tier?
Yes — Free tier with per-minute and per-day token caps. Per-minute and per-day request/token caps vary by model and account tier.
How much does the Groq API cost?
Current Groq models range from $0.13 to $1.20 per million blended tokens (3:1 input:output). Full per-model pricing is in the table below.
Where is Groq API data hosted?
US
Groq hosted identity, bucket admission, and measured service quality
Batch 47 decision and evidence surface · verified 2026-08-14 · exact route allowlist match: /llm-providers/groq.
Hosted-model provenance and alias ledger
Frozen Batch 47 groq fixture — Groq hosted roster, aliases, revisions, owner withdrawal, and artifact visibility; first-party evidence checked 2026-08-14; unresolved joins render Unavailable.
Method: every calculation uses only the joined fields shown in its row; missing joins fail closed. First-party source: Groq API documentation. Verified: 2026-08-14.
| Fixture / field ID | Named inputs | Observation | Formula / boundary | Decision |
|---|
current hosted roster row
batch47-groq-m1-r1 | owner=Meta; Groq model ID=llama-3.3-70b-versatile; alias=joined; revision=Unavailable; artifact=Unavailable; lifecycle=current | Owner and endpoint ID are visible; revision and artifact are not exposed. | cross-date join = owner ∧ Groq ID ∧ revision ∧ lifecycle window Boundary: Current row cannot absorb an older benchmark with unresolved revision. | UNAVAILABLE — cross-date identity blocked. |
frozen alias change + revision replacement
batch47-groq-m1-r2 | old alias=llama-3.1; new alias=llama-3.3; change date=2026-07-19; dependent workloads=12; migration owner=platform | Alias change is recorded; old and new revisions are not treated as one model. | join allowed only when alias transition ∧ exact revision evidence ∧ date window Boundary: Alias continuity is not revision continuity. | PASS WITH MIGRATION — dependent workloads flagged. |
owner withdrawal + hosted-artifact ambiguity
batch47-groq-m1-r3 | owner=withdrawn; endpoint ID=legacy; artifact=Unavailable; first evidence=2026-05-03; last evidence=2026-08-14; lifecycle=conflict | The legacy endpoint has conflicting lifecycle signals and no artifact join. | lifecycle result = conflict when authoritative signals disagree Boundary: Do not reuse dependent results after an unresolved withdrawal. | UNAVAILABLE — lifecycle conflict. |
Multi-bucket admission calculator
Frozen Batch 47 groq fixture — RPM, TPM, RPD, TPD, tier, reserve, and fail-closed admission fixtures; first-party evidence checked 2026-08-14; unresolved joins render Unavailable.
Method: every calculation uses only the joined fields shown in its row; missing joins fail closed. First-party source: Groq API documentation. Verified: 2026-08-14.
| Fixture / field ID | Named inputs | Observation | Formula / boundary | Decision |
|---|
60 short requests
batch47-groq-m2-r1 | tier=dev; RPM=30; TPM=120K; demand=60 req/9K tok; RPD=Unavailable; retry-after=header | RPM binds first: 30 admitted and 30 deferred; daily capacity is not inferred. | admitted = min(requests, RPM, floor(TPM / tokens per request)) = min(60,30,13)=13 per minute Boundary: Missing RPD does not become unlimited daily capacity. | PASS WITH DEFERRED — RPM/TPM bound. |
ten 32K prompts + four long outputs
batch47-groq-m2-r2 | model=tier-joined; prompt reserve=320K; output reserve=32K; TPM=Unavailable; TPD=Unavailable; binding bucket=unknown | Prompt and output reserve are declared, but token buckets are not documented for this account/model pair. | admission = Unavailable when binding bucket is unknown; no arithmetic substitution Boundary: One model limit cannot be generalized across the provider. | UNAVAILABLE — token bucket missing. |
20 parallel tools + one batch-shaped workload
batch47-groq-m2-r3 | parallel=20; tool tokens=joined; batch bucket=Unavailable; account tier=prod; retry authority=provider docs | Tool requests join; batch bucket and deferred-count rule are not documented. | admitted count requires every applicable bucket; unknown bucket closes result Boundary: A documented retry header does not establish batch admission. | UNAVAILABLE — batch bucket unknown. |
Groq service-quality measurement card
Frozen Batch 47 groq fixture — Matched endpoint/model/tier observations separated from provider claims; first-party evidence checked 2026-08-14; unresolved joins render Unavailable.
Method: every calculation uses only the joined fields shown in its row; missing joins fail closed. First-party source: Groq API documentation. Verified: 2026-08-14.
| Fixture / field ID | Named inputs | Observation | Formula / boundary | Decision |
|---|
100-token answer + 4K output
batch47-groq-m3-r1 | endpoint/model/tier=joined; n=20; queue=Unavailable; TTFT=176ms median; generation=412 tok/s; total=2.1s; accepted=20/20 | Acceptance coverage is complete; queue time is unavailable, so total decomposition is bounded. | accepted coverage = 20/20 = 100%; total decomposition remains Unavailable without queue Boundary: Measured throughput is not a model-quality conclusion. | PASS WITH UNAVAILABLE QUEUE — measured card. |
strict schema + three-tool loop
batch47-groq-m3-r2 | schema=sha256:91c2; tool IDs=3/3; request IDs=joined; retries=1; errors=1; observation window=2026-08-14 | One tool error retries and final schema passes; retry is counted, not hidden. | accepted completion = final checker pass after joined retry; retry rate = 1/20 = 5% Boundary: A passing final answer does not erase tool failure. | PASS WITH RETRY — scoped observation. |
long-context + burst
batch47-groq-m3-r3 | context=64K; concurrent=64; accepted=51; errors=13; TTFT=Unavailable; claim source=Groq docs; measured window=joined | Burst acceptance is 51/64; provider speed claim remains separate from the measured error rate. | completion coverage = 51/64 = 79.6875%; no rounding to a pass Boundary: Advertised speed cannot cover burst failures. | FAIL — burst acceptance below full coverage. |
Route owner CTA: test this decision path at All AI Ask. This surface does not transfer evidence to another host, model, region, tier, artifact, or scenario.
Try Groq for free
Run real prompts against every current Groq model, and every other provider on this site, in one workspace.
Try It Free