Cerebras GLM-4.7 API Pricing: Frontier Bilingual Intelligence at 1,500+ TPS
Examine Cerebras GLM-4.7 API pricing ($2.25/M input, $2.75/M output), extreme bilingual inference speed on Cerebras CS-3, and enterprise reasoning throughput.
Full specs, context window and API limits →How much does GLM 4.7 (Cerebras) cost per million tokens?
Cerebras GLM-4.7 costs $2.25 per million input tokens and $2.75 per million output tokens ($2.375/M blended at 3:1). Combines premier Chinese/English reasoning with Cerebras CS-3 wafer-scale throughput. Verified 2026-09-08.
How much does GLM 4.7 (Cerebras) cost per 1,000 requests?
Computed from generated token pricing. Each row assumes the listed input and output tokens per request; output is adjusted by this model's measured 7.53× verbosity factor.
| Request shape | Input tokens | Output tokens | Cost / 1,000 requests |
|---|---|---|---|
| Short | 100 | 50 | $1.2604 |
| Medium | 1,000 | 500 | $12.6038 |
| Long | 4,000 | 2,000 | $50.4150 |
Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Verbosity run: 2026-06-16T20:31:30.728Z.
Three model-specific pricing decisions
This is the Cerebras-hosted GLM 4.7 owner. GLM-5.2 direct-API economics are a dated comparison input; provider policy and open-weight deployment remain separate.
1. Coding-agent and long-output Cerebras bills
| Workload | Input / output | 100K bill | Cerebras delivery evidence |
|---|---|---|---|
| Coding agent | 8,000 / 1,200 | $2130.00 | 1980 tokens/sec; TTFT 110 ms; 5 measured samples |
| Long output | 32,000 / 4,000 | $8300.00 | 1980 tokens/sec; TTFT 110 ms; 5 measured samples |
2. Required quality/retry uplift versus GLM-5.2 direct
| Fixed input / output | GLM 4.7 (Cerebras) | GLM-5.2 | Narrow decision boundary |
|---|---|---|---|
| Coding agent · 8,000 / 1,200 | $2130.00 | $1648.00 | 15% accepted-result uplift required |
| Long output · 32,000 / 4,000 | $8300.00 | $6240.00 | 20% accepted-result uplift required |
Formula: requests × (input tokens × input $/M + output tokens × output $/M) ÷ 1,000,000. The uplift is a planning threshold, not a measured quality claim.
3. Capacity and deployment-risk ledger
| Dimension | Cerebras/API evidence | Direct comparison boundary |
|---|---|---|
| Token spend | $2.25 / $2.75 per M | Sourced |
| Completion time | 1980 tokens/sec; TTFT 110 ms; 5 measured samples | Output-only estimate excludes queueing |
| Context fit / quota / region / currency | Unavailable | Do not infer |
| Cache / batch / SLA | Unavailable | No free or guaranteed behavior |
Verified 2026-06-14. Luna is the data owner. “Unavailable” means no compatible dated evidence was found; it is never treated as zero. First-party price source · Run this Batch 5 scenario.
All three Batch 5 decisions are server-rendered for GLM 4.7 (Cerebras); fixed inputs, formulas, dated sources, speed sample state, and unavailable mechanics are visible.
Batch 68 · exact model pricing decision contributions · verified 2026-09-08
Exact model boundary: Cerebras cerebras-glm-4.7 (slug cerebras-glm-4-7). First-party provider pricing and API documentation remain fact owners.
Wafer-scale bilingual token pricing and monthly spend matrix
Frozen Batch 68 scenario board. Formula / deterministic rule: monthly_spend = calls * ((in_tokens * 2.25 + out_tokens * 2.75) / 1M) Boundary: Owns Cerebras CS-3 wafer-scale bilingual token tariff modeling.
| Frozen scenario / field ID | Model, identity, provider, and evidence fields | Result | State |
|---|---|---|---|
batch68-cerebras-glm-4-7-m1-r120K complex bilingual contract reviews | model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=20K complex bilingual contract reviews; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — 20K complex bilingual contract reviews is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch68-cerebras-glm-4-7-m1-r2100K cross-border enterprise inquiries | model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=100K cross-border enterprise inquiries; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — 100K cross-border enterprise inquiries is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch68-cerebras-glm-4-7-m1-r3500K automated live-data monitoring passes | model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=500K automated live-data monitoring passes; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — 500K automated live-data monitoring passes is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch68-cerebras-glm-4-7-m1-r4batch offline ingestion queue | model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=batch offline ingestion queue; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — batch offline ingestion queue is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch68-cerebras-glm-4-7-m1-r5enterprise dedicated CS-3 provisioning | model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=enterprise dedicated CS-3 provisioning; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — enterprise dedicated CS-3 provisioning is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch68-cerebras-glm-4-7-m1-r6unresolved billing currency | model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=unresolved billing currency; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — unresolved billing currency has no matched, dated bilateral observation. | FAIL CLOSED — manual, probe, or source evidence required |
First-party provenance: Cerebras official pricing; verification date 2026-09-08. Missing or conflicting joins fail closed.
Extreme bilingual inference SLA and turnaround audit
Frozen Batch 68 scenario board. Formula / deterministic rule: turnaround = ttft + (tokens_out / tps); Cerebras CS-3 wafer-scale engine Boundary: Owns turnaround time benchmarks and bilingual user experience responsiveness.
| Frozen scenario / field ID | Model, identity, provider, and evidence fields | Result | State |
|---|---|---|---|
batch68-cerebras-glm-4-7-m2-r1sub-80ms bilingual streaming dialogue | model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=sub-80ms bilingual streaming dialogue; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — sub-80ms bilingual streaming dialogue is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch68-cerebras-glm-4-7-m2-r2complex Chinese mathematical proof (<200ms) | model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=complex Chinese mathematical proof (<200ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — complex Chinese mathematical proof (<200ms) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch68-cerebras-glm-4-7-m2-r3instant multi-page legal translation (<400ms) | model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=instant multi-page legal translation (<400ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — instant multi-page legal translation (<400ms) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch68-cerebras-glm-4-7-m2-r4high-concurrency request surge | model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=high-concurrency request surge; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — high-concurrency request surge is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch68-cerebras-glm-4-7-m2-r5network transit buffer | model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=network transit buffer; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — network transit buffer is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch68-cerebras-glm-4-7-m2-r6unmeasured speed fixture | model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=unmeasured speed fixture; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — unmeasured speed fixture is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
First-party provenance: Cerebras documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
Dedicated CS-3 wafer cluster vs managed cloud API break-even
Frozen Batch 68 scenario board. Formula / deterministic rule: managed_cost = tokens * blended_rate; cluster_cost = dedicated_hardware_monthly Boundary: Owns private dedicated wafer-scale deployment versus serverless API billing.
| Frozen scenario / field ID | Model, identity, provider, and evidence fields | Result | State |
|---|---|---|---|
batch68-cerebras-glm-4-7-m3-r1intermittent corporate workload (<10M tokens/mo) | model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=intermittent corporate workload (<10M tokens/mo); monthly token volume; managed Cerebras spend; dedicated CS-3 monthly TCO; financial break-even; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — intermittent corporate workload (<10M tokens/mo) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch68-cerebras-glm-4-7-m3-r250M monthly tokens (API optimal) | model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=50M monthly tokens (API optimal); monthly token volume; managed Cerebras spend; dedicated CS-3 monthly TCO; financial break-even; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — 50M monthly tokens (API optimal) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch68-cerebras-glm-4-7-m3-r3250M high-volume enterprise scale | model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=250M high-volume enterprise scale; monthly token volume; managed Cerebras spend; dedicated CS-3 monthly TCO; financial break-even; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — 250M high-volume enterprise scale is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch68-cerebras-glm-4-7-m3-r41B+ constant saturation (dedicated threshold) | model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=1B+ constant saturation (dedicated threshold); monthly token volume; managed Cerebras spend; dedicated CS-3 monthly TCO; financial break-even; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — 1B+ constant saturation (dedicated threshold) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch68-cerebras-glm-4-7-m3-r5uncommitted hardware idle hours | model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=uncommitted hardware idle hours; monthly token volume; managed Cerebras spend; dedicated CS-3 monthly TCO; financial break-even; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — uncommitted hardware idle hours is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch68-cerebras-glm-4-7-m3-r6unresolved data center energy tariff | model=cerebras-glm-4.7; slug=cerebras-glm-4-7; provider=Cerebras; scenario=unresolved data center energy tariff; monthly token volume; managed Cerebras spend; dedicated CS-3 monthly TCO; financial break-even; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — unresolved data center energy tariff has no matched, dated bilateral observation. | FAIL CLOSED — manual, probe, or source evidence required |
First-party provenance: Cerebras official pricing; verification date 2026-09-08. Missing or conflicting joins fail closed.
Method and limitations: formulas are deterministic; observed and assumed inputs are labeled; no missing provider, host, account, region, realm, alias, snapshot, revision, weight, artifact, control, tool, modality, workload, rate-period, timestamp, or result is transferred. Run the cerebras-glm-4-7 Batch 68 scenario →
How fast is GLM 4.7 (Cerebras)?
How much does GLM 4.7 (Cerebras) cost at scale?
| Tokens / month | Est. cost (blended 3:1) |
|---|---|
| 100,000 | $0.24 |
| 1,000,000 | $2.38 |
| 10,000,000 | $23.75 |
| 100,000,000 | $237.50 |
How does GLM 4.7 (Cerebras) compare with other models?
What is GLM 4.7 (Cerebras) best for?
What should you explore next for GLM 4.7 (Cerebras)?
What are common questions about GLM 4.7 (Cerebras)?
Is GLM 4.7 (Cerebras) cheaper than GPT-5.6 Luna?
GLM 4.7 (Cerebras) costs $2.38/M blended tokens, GPT-5.6 Luna costs $2.25/M — GPT-5.6 Luna is cheaper.
How much does 1 million tokens cost with GLM 4.7 (Cerebras)?
At a 3:1 input:output ratio, 1 million blended tokens costs approximately $2.38. Pure input costs $2.25/M; pure output costs $2.75/M.
What does GLM 4.7 (Cerebras) cost at high volume?
At 100 million blended tokens a month, GLM 4.7 (Cerebras) costs approximately $237.50. See the cost-at-scale table below for other volumes.
