← Back to all pricing

Cerebras GLM-4.7 API Pricing: Frontier Bilingual Intelligence at 1,500+ TPS

Examine Cerebras GLM-4.7 API pricing ($2.25/M input, $2.75/M output), extreme bilingual inference speed on Cerebras CS-3, and enterprise reasoning throughput.

Full specs, context window and API limits →

How much does GLM 4.7 (Cerebras) cost per million tokens?

Cerebras GLM-4.7 costs $2.25 per million input tokens and $2.75 per million output tokens ($2.375/M blended at 3:1). Combines premier Chinese/English reasoning with Cerebras CS-3 wafer-scale throughput. Verified 2026-09-08.

Verified 2026-09-08 source
Input
$2.25/M
Output
$2.75/M
Blended
$2.38/M
Provider
Verified 2026-06-14source

How much does GLM 4.7 (Cerebras) cost per 1,000 requests?

Computed from generated token pricing. Each row assumes the listed input and output tokens per request; output is adjusted by this model's measured 7.53× verbosity factor.

Request shapeInput tokensOutput tokensCost / 1,000 requests
Short10050$1.2604
Medium1,000500$12.6038
Long4,0002,000$50.4150

Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Verbosity run: 2026-06-16T20:31:30.728Z.

Three model-specific pricing decisions

This is the Cerebras-hosted GLM 4.7 owner. GLM-5.2 direct-API economics are a dated comparison input; provider policy and open-weight deployment remain separate.

1. Coding-agent and long-output Cerebras bills

WorkloadInput / output100K billCerebras delivery evidence
Coding agent8,000 / 1,200$2130.001980 tokens/sec; TTFT 110 ms; 5 measured samples
Long output32,000 / 4,000$8300.001980 tokens/sec; TTFT 110 ms; 5 measured samples

2. Required quality/retry uplift versus GLM-5.2 direct

Fixed input / outputGLM 4.7 (Cerebras)GLM-5.2Narrow decision boundary
Coding agent · 8,000 / 1,200$2130.00$1648.0015% accepted-result uplift required
Long output · 32,000 / 4,000$8300.00$6240.0020% accepted-result uplift required

Formula: requests × (input tokens × input $/M + output tokens × output $/M) ÷ 1,000,000. The uplift is a planning threshold, not a measured quality claim.

3. Capacity and deployment-risk ledger

DimensionCerebras/API evidenceDirect comparison boundary
Token spend$2.25 / $2.75 per MSourced
Completion time1980 tokens/sec; TTFT 110 ms; 5 measured samplesOutput-only estimate excludes queueing
Context fit / quota / region / currencyUnavailableDo not infer
Cache / batch / SLAUnavailableNo free or guaranteed behavior

Verified 2026-06-14. Luna is the data owner. “Unavailable” means no compatible dated evidence was found; it is never treated as zero. First-party price source · Run this Batch 5 scenario.

All three Batch 5 decisions are server-rendered for GLM 4.7 (Cerebras); fixed inputs, formulas, dated sources, speed sample state, and unavailable mechanics are visible.

Batch 68 · exact model pricing decision contributions · verified 2026-09-08

Exact model boundary: Cerebras cerebras-glm-4.7 (slug cerebras-glm-4-7). First-party provider pricing and API documentation remain fact owners.

Wafer-scale bilingual token pricing and monthly spend matrix

Frozen Batch 68 scenario board. Formula / deterministic rule: monthly_spend = calls * ((in_tokens * 2.25 + out_tokens * 2.75) / 1M) Boundary: Owns Cerebras CS-3 wafer-scale bilingual token tariff modeling.

Frozen scenario / field IDModel, identity, provider, and evidence fieldsResultState
batch68-cerebras-glm-4-7-m1-r1
20K complex bilingual contract reviews
model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=20K complex bilingual contract reviews; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 20K complex bilingual contract reviews is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-glm-4-7-m1-r2
100K cross-border enterprise inquiries
model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=100K cross-border enterprise inquiries; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 100K cross-border enterprise inquiries is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-glm-4-7-m1-r3
500K automated live-data monitoring passes
model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=500K automated live-data monitoring passes; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 500K automated live-data monitoring passes is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-glm-4-7-m1-r4
batch offline ingestion queue
model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=batch offline ingestion queue; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — batch offline ingestion queue is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-glm-4-7-m1-r5
enterprise dedicated CS-3 provisioning
model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=enterprise dedicated CS-3 provisioning; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — enterprise dedicated CS-3 provisioning is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-glm-4-7-m1-r6
unresolved billing currency
model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=unresolved billing currency; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — unresolved billing currency has no matched, dated bilateral observation.FAIL CLOSED — manual, probe, or source evidence required

First-party provenance: Cerebras official pricing; verification date 2026-09-08. Missing or conflicting joins fail closed.

Extreme bilingual inference SLA and turnaround audit

Frozen Batch 68 scenario board. Formula / deterministic rule: turnaround = ttft + (tokens_out / tps); Cerebras CS-3 wafer-scale engine Boundary: Owns turnaround time benchmarks and bilingual user experience responsiveness.

Frozen scenario / field IDModel, identity, provider, and evidence fieldsResultState
batch68-cerebras-glm-4-7-m2-r1
sub-80ms bilingual streaming dialogue
model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=sub-80ms bilingual streaming dialogue; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — sub-80ms bilingual streaming dialogue is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-glm-4-7-m2-r2
complex Chinese mathematical proof (<200ms)
model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=complex Chinese mathematical proof (<200ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — complex Chinese mathematical proof (<200ms) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-glm-4-7-m2-r3
instant multi-page legal translation (<400ms)
model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=instant multi-page legal translation (<400ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — instant multi-page legal translation (<400ms) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-glm-4-7-m2-r4
high-concurrency request surge
model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=high-concurrency request surge; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — high-concurrency request surge is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-glm-4-7-m2-r5
network transit buffer
model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=network transit buffer; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — network transit buffer is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-glm-4-7-m2-r6
unmeasured speed fixture
model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=unmeasured speed fixture; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — unmeasured speed fixture is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict

First-party provenance: Cerebras documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

Dedicated CS-3 wafer cluster vs managed cloud API break-even

Frozen Batch 68 scenario board. Formula / deterministic rule: managed_cost = tokens * blended_rate; cluster_cost = dedicated_hardware_monthly Boundary: Owns private dedicated wafer-scale deployment versus serverless API billing.

Frozen scenario / field IDModel, identity, provider, and evidence fieldsResultState
batch68-cerebras-glm-4-7-m3-r1
intermittent corporate workload (<10M tokens/mo)
model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=intermittent corporate workload (<10M tokens/mo); monthly token volume; managed Cerebras spend; dedicated CS-3 monthly TCO; financial break-even; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — intermittent corporate workload (<10M tokens/mo) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-glm-4-7-m3-r2
50M monthly tokens (API optimal)
model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=50M monthly tokens (API optimal); monthly token volume; managed Cerebras spend; dedicated CS-3 monthly TCO; financial break-even; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 50M monthly tokens (API optimal) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-glm-4-7-m3-r3
250M high-volume enterprise scale
model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=250M high-volume enterprise scale; monthly token volume; managed Cerebras spend; dedicated CS-3 monthly TCO; financial break-even; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 250M high-volume enterprise scale is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-glm-4-7-m3-r4
1B+ constant saturation (dedicated threshold)
model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=1B+ constant saturation (dedicated threshold); monthly token volume; managed Cerebras spend; dedicated CS-3 monthly TCO; financial break-even; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 1B+ constant saturation (dedicated threshold) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-glm-4-7-m3-r5
uncommitted hardware idle hours
model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=uncommitted hardware idle hours; monthly token volume; managed Cerebras spend; dedicated CS-3 monthly TCO; financial break-even; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — uncommitted hardware idle hours is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-glm-4-7-m3-r6
unresolved data center energy tariff
model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=unresolved data center energy tariff; monthly token volume; managed Cerebras spend; dedicated CS-3 monthly TCO; financial break-even; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — unresolved data center energy tariff has no matched, dated bilateral observation.FAIL CLOSED — manual, probe, or source evidence required

First-party provenance: Cerebras official pricing; verification date 2026-09-08. Missing or conflicting joins fail closed.

Method and limitations: formulas are deterministic; observed and assumed inputs are labeled; no missing provider, host, account, region, realm, alias, snapshot, revision, weight, artifact, control, tool, modality, workload, rate-period, timestamp, or result is transferred. Run the cerebras-glm-4-7 Batch 68 scenario →

How fast is GLM 4.7 (Cerebras)?

Tokens / sec
1980
TTFT
110 ms
Rank
#2 of 31
$ / M ÷ t/s
$0.0012
Measured with 5 runs on a fixed prompt — see the full methodology.

How much does GLM 4.7 (Cerebras) cost at scale?

Tokens / monthEst. cost (blended 3:1)
100,000$0.24
1,000,000$2.38
10,000,000$23.75
100,000,000$237.50

How does GLM 4.7 (Cerebras) compare with other models?

GPT-OSS 120B (Cerebras)$0.45/MGPT-5.6 Luna$2.25/MGrok-3$2.50/MGLM-5.2$2.15/M
See all Cerebras models →

What is GLM 4.7 (Cerebras) best for?

#5 for Agents & Tool Use#7 for Chatbots & Support#8 for Translation
Looking for a cheaper option?
GPT-OSS 20B is 94.5% cheaper — a config migration. See all 8 alternatives to GLM 4.7 (Cerebras)

What are common questions about GLM 4.7 (Cerebras)?

Is GLM 4.7 (Cerebras) cheaper than GPT-5.6 Luna?

GLM 4.7 (Cerebras) costs $2.38/M blended tokens, GPT-5.6 Luna costs $2.25/M — GPT-5.6 Luna is cheaper.

How much does 1 million tokens cost with GLM 4.7 (Cerebras)?

At a 3:1 input:output ratio, 1 million blended tokens costs approximately $2.38. Pure input costs $2.25/M; pure output costs $2.75/M.

What does GLM 4.7 (Cerebras) cost at high volume?

At 100 million blended tokens a month, GLM 4.7 (Cerebras) costs approximately $237.50. See the cost-at-scale table below for other volumes.

Try GLM 4.7 (Cerebras) for free

Run real prompts against GLM 4.7 (Cerebras) and every other model on this page in one workspace.

Try GLM 4.7 (Cerebras) Free