Cerebras Rate Limits by Tier

What are Cerebras's API rate limits?

Cerebras's entry tier (Free) allows an unpublished number of requests/min and an unpublished number of tokens/min for documented default/model family. Limits scale up through Production as cumulative spend and account age increase — see the full table below, verified 2026-08-14.

Verified 2026-08-14 source

Limits by tier

TierQualificationModel classRPMTPMRPDConcurrent
FreeFree account quotadocumented default/model family
ProductionAccount quota shown in consoledocumented default/model family

— means not documented by Cerebras, never a guess.

What this means for your workload

Classification at volume: 116 calls/min and 60,320 tokens/min at the production profile.

This provider publishes no numeric cap for this workload; check the account console before launch.

Response headers

retry-afterSeconds to wait before retrying, when supplied with a 429
rate-limit response headersProvider-specific remaining and reset counters when documented

When you exceed the limit

Cerebras returns HTTP 429.

Reduce concurrency and consult the account quota.

FAQ

What happens when I exceed Cerebras's rate limit?

Cerebras returns HTTP 429. Reduce concurrency and consult the account quota.

How do I request a rate limit increase on Cerebras?

Request one from the account dashboard: https://cloud.cerebras.ai

Cerebras provider hubGet a Cerebras API key

Batch 48 · cerebras rate-limit decision and evidence modules. Surface verification: 2026-08-14. These are dated, route-local fixtures, not live quota claims.

Cerebras organization-project ceiling resolver

Frozen cerebras evidence board: one project · two projects under one organization · project below/above organization · free/developer · enterprise override. Formula / decision rule: effective limit = min(project, organization) where both are known

Frozen fixtureInputs and observationFormula / boundaryState
batch48-cerebras-m1-r1
one project · two projects under one organization
principal/account=joined; endpoint=joined; model/pool=joined; bucket/window=joined; observed=2026-08-14
The primary identity and product joins are present for this frozen fixture.
effective limit = min(project, organization) where both are known
A missing provider, product, account, endpoint, model, bucket, or time-window join fails closed; no value is copied across routes.
PASS WITH SCOPE — frozen observation only.
batch48-cerebras-m1-r2
project below/above organization
request/response/error=joined; limit/remaining/reset=observed; retry identity=joined; region=joined
The response evidence is usable only for the named bucket and time window; no neighboring provider is borrowed.
effective limit = min(project, organization) where both are known
A missing provider, product, account, endpoint, model, bucket, or time-window join fails closed; no value is copied across routes.
PASS WITH LIMIT — dated evidence only.
batch48-cerebras-m1-r3
free/developer · enterprise override
workload=joined; prompt/output reserve=explicit; shared pool=declared; unknown=Unavailable; owner=platform
Admission or recovery is withheld where the account-specific value or side-effect identity is absent.
effective limit = min(project, organization) where both are known
A missing provider, product, account, endpoint, model, bucket, or time-window join fails closed; no value is copied across routes.
UNAVAILABLE — live owner evidence required.

Provenance: Frozen Batch 48 cerebras fixture: module 1 observation and calculation. First-party evidence checked 2026-08-14; unresolved joins fail closed. Surface verification date: 2026-08-14. Cerebras Inference docs (first-party source)

Cerebras rate-header and reset ledger

Frozen cerebras evidence board: normal response · request-day near-limit · token-minute near-limit · simultaneous exhaustion · 429 · missing header · malformed reset. Formula / decision rule: reset timestamp = observation time + provider reset unit; malformed or missing reset → Unavailable

Frozen fixtureInputs and observationFormula / boundaryState
batch48-cerebras-m2-r1
normal response · request-day near-limit
principal/account=joined; endpoint=joined; model/pool=joined; bucket/window=joined; observed=2026-08-14
The primary identity and product joins are present for this frozen fixture.
reset timestamp = observation time + provider reset unit; malformed or missing reset → Unavailable
A missing provider, product, account, endpoint, model, bucket, or time-window join fails closed; no value is copied across routes.
PASS WITH SCOPE — frozen observation only.
batch48-cerebras-m2-r2
token-minute near-limit · simultaneous exhaustion
request/response/error=joined; limit/remaining/reset=observed; retry identity=joined; region=joined
The response evidence is usable only for the named bucket and time window; no neighboring provider is borrowed.
reset timestamp = observation time + provider reset unit; malformed or missing reset → Unavailable
A missing provider, product, account, endpoint, model, bucket, or time-window join fails closed; no value is copied across routes.
PASS WITH LIMIT — dated evidence only.
batch48-cerebras-m2-r3
429 · missing header · malformed reset
workload=joined; prompt/output reserve=explicit; shared pool=declared; unknown=Unavailable; owner=platform
Admission or recovery is withheld where the account-specific value or side-effect identity is absent.
reset timestamp = observation time + provider reset unit; malformed or missing reset → Unavailable
A missing provider, product, account, endpoint, model, bucket, or time-window join fails closed; no value is copied across routes.
UNAVAILABLE — live owner evidence required.

Provenance: Frozen Batch 48 cerebras fixture: module 2 observation and calculation. First-party evidence checked 2026-08-14; unresolved joins fail closed. Surface verification date: 2026-08-14. Cerebras Inference docs (first-party source)

Cerebras burst-and-long-context admission board

Frozen cerebras evidence board: 100 short chats · eight long prompts · four long outputs · 16 strict extractions · 64-request burst · retry wave. Formula / decision rule: completion coverage = admitted / declared; inference speed is not quota capacity

Frozen fixtureInputs and observationFormula / boundaryState
batch48-cerebras-m3-r1
100 short chats · eight long prompts
principal/account=joined; endpoint=joined; model/pool=joined; bucket/window=joined; observed=2026-08-14
The primary identity and product joins are present for this frozen fixture.
completion coverage = admitted / declared; inference speed is not quota capacity
A missing provider, product, account, endpoint, model, bucket, or time-window join fails closed; no value is copied across routes.
PASS WITH SCOPE — frozen observation only.
batch48-cerebras-m3-r2
four long outputs · 16 strict extractions
request/response/error=joined; limit/remaining/reset=observed; retry identity=joined; region=joined
The response evidence is usable only for the named bucket and time window; no neighboring provider is borrowed.
completion coverage = admitted / declared; inference speed is not quota capacity
A missing provider, product, account, endpoint, model, bucket, or time-window join fails closed; no value is copied across routes.
PASS WITH LIMIT — dated evidence only.
batch48-cerebras-m3-r3
64-request burst · retry wave
workload=joined; prompt/output reserve=explicit; shared pool=declared; unknown=Unavailable; owner=platform
Admission or recovery is withheld where the account-specific value or side-effect identity is absent.
completion coverage = admitted / declared; inference speed is not quota capacity
A missing provider, product, account, endpoint, model, bucket, or time-window join fails closed; no value is copied across routes.
UNAVAILABLE — live owner evidence required.

Provenance: Frozen Batch 48 cerebras fixture: module 3 observation and calculation. First-party evidence checked 2026-08-14; unresolved joins fail closed. Surface verification date: 2026-08-14. Cerebras Inference docs (first-party source)

Run the cerebras Batch 48 evidence scenario →