Z.ai Rate Limits by Tier

What are Z.ai's API rate limits?

Z.ai's entry tier (Default) allows an unpublished number of requests/min and an unpublished number of tokens/min for documented default/model family. Limits scale up through Default as cumulative spend and account age increase — see the full table below, verified 2026-08-14.

Verified 2026-08-14 source

Limits by tier

TierQualificationModel classRPMTPMRPDConcurrent
DefaultAccount quota shown in the Z.ai consoledocumented default/model family

— means not documented by Z.ai, never a guess.

What this means for your workload

Classification at volume: 116 calls/min and 60,320 tokens/min at the production profile.

This provider publishes no numeric cap for this workload; check the account console before launch.

Response headers

retry-afterSeconds to wait before retrying, when supplied with a 429
rate-limit response headersProvider-specific remaining and reset counters when documented

When you exceed the limit

Z.ai returns HTTP 429.

Back off and check the model-specific account quota.

FAQ

What happens when I exceed Z.ai's rate limit?

Z.ai returns HTTP 429. Back off and check the model-specific account quota.

How do I request a rate limit increase on Z.ai?

Request one from the account dashboard: https://bigmodel.cn/console

Z.ai provider hubGet a Z.ai API key

Batch 49 · zai decision and evidence contributions. Surface verification: 2026-08-14. These are route-local, server-rendered fixtures; unavailable values are not inferred.

Z.ai observed-limit snapshot ledger

Frozen Batch 49 fixture board. Formula / decision rule: effective value = exact account + key + endpoint + model + plan + window observation Boundary: A sibling key, alias, or plan cannot supply an undocumented value.

Frozen fixture / field IDJoined inputs and observationCalculated resultState
batch49-zai-m1-r1
general API key · exact GLM alias
account=acct-zai-12; key=fp-zai-g-01; endpoint=api.z.ai; model=GLM-5; plan=general; dimension=RPM; window=minute; console=observed 60
The console value is joined to the account, key class, endpoint, exact alias, and minute window.
RPM=60; sharing edge=account-level only where console states itPASS WITH SCOPE — observed limit.
batch49-zai-m1-r2
second key under one account · async job
account=acct-zai-12; key=fp-zai-g-02; endpoint=api.z.ai; job=async-14; dimension=concurrency; value=Unavailable; source=response headers
The second key joins the account, but no async concurrency ceiling is returned.
concurrency = Unavailable; no copy from general RPMUNAVAILABLE — live account limit required.
batch49-zai-m1-r3
coding-compatible endpoint · enterprise override
account=acct-zai-99; endpoint=code.z.ai; model=GLM-Coder; plan=enterprise; override=undocumented; observed-at=2026-08-14
An operator reports an override, but no console or response evidence joins it.
enterprise value = Unavailable; report only the observation provenanceFAIL CLOSED — undocumented override.

Provenance: Batch 49 zai module 1 first-party evidence and surface verification date 2026-08-14. Z.ai developer limits documentation. Missing joins fail closed.

Z.ai 429 evidence-and-replay classifier

Frozen Batch 49 fixture board. Formula / decision rule: replay = provider evidence ∧ debit visibility ∧ idempotency/effect key Boundary: Backoff changes timing; it cannot make a non-idempotent replay safe.

Frozen fixture / field IDJoined inputs and observationCalculated resultState
batch49-zai-m2-r1
request-frequency / token-budget 429
HTTP=429; code=rate_limit; request=req-zai-21; tool=none; reset=5s; debit=visible; idempotency=key-z-1
Returned timing and debit evidence identify a transient request/token class.
retry after 5s with bounded jitter; replay=eligibleRETRY ELIGIBLE — identity joined.
batch49-zai-m2-r2
partial stream · tool side effect
HTTP=429; request=req-zai-22; stream=partial; tool=charge-card; tool_id=tool-88; effect checkpoint=missing
A side-effecting tool call lacks a checkpoint, so replay could duplicate the effect.
retry=blocked regardless of exponential backoffFAIL CLOSED — operator reconciliation.
batch49-zai-m2-r3
opaque 429 · insufficient resource or overload
HTTP=429; code=opaque; request=req-zai-23; reset=missing; debit=missing; concurrency=joined
The response cannot distinguish resource entitlement from transient overload.
class={resource, overload}; next action=collect console/support evidenceUNRESOLVED — no automatic replay.

Provenance: Batch 49 zai module 2 first-party evidence and surface verification date 2026-08-14. Z.ai developer limits documentation. Missing joins fail closed.

GLM workload admission board

Frozen Batch 49 fixture board. Formula / decision rule: capacity verdict requires a live account limit for every applicable request/token/concurrency bucket Boundary: Unknown capacity is neither zero nor unlimited.

Frozen fixture / field IDJoined inputs and observationCalculated resultState
batch49-zai-m3-r1
100 short calls
account=acct-zai-12; key=fp-zai-g-01; request reserve=100; token reserve=20000; RPM=60; TPM=Unavailable
Request bucket alone admits only the first 60 calls; token bucket is unknown.
admitted=60; deferred=40; unknown token impactDEFER — token ceiling missing.
batch49-zai-m3-r2
ten 32K prompts · four long outputs · 20 parallel tools
prompt=320000; output=64000; tools=20; concurrency=Unavailable; model=GLM-5
Exact model joins, but concurrency and token reserves do not.
admitted=0; deferred=34; binding bucket=UnavailableUNAVAILABLE — live limits required.
batch49-zai-m3-r3
asynchronous batch · retry wave
job=async-19; retry_count=3; shared_pool=account; idempotency=partial; reset=Unavailable
The account pool is known, but reset timing and all retry identities are not.
next-safe schedule = Unavailable; do not enqueue retry waveFAIL CLOSED — incomplete replay joins.

Provenance: Batch 49 zai module 3 first-party evidence and surface verification date 2026-08-14. Z.ai developer limits documentation. Missing joins fail closed.

Run the zai Batch 49 evidence scenario →