GLM 4.7 (Cerebras)
Coding agents that need GLM-class quality with Cerebras-level speed.
What are GLM 4.7 (Cerebras)'s specs and price?
GLM 4.7 (Cerebras), built by Cerebras, ships a 200K-token context window and a 33K-token max output, released 2026-01. It supports text input with a dedicated reasoning mode and costs $2.38 per million blended tokens, the 23rd-cheapest of 39 models we track.
Batch 43 evidence surface · verified 2026-08-27 · frozen route allowlist: /models/cerebras-glm-4-7
Cerebras GLM 4.7 identity, migration, and coding-agent frontier
Batch 43 · M1: Cerebras GLM identity ledger
Formula: Identity pass = exact zai-glm-4.7 record ∧ Z.ai revision ∧ public variant ∧ host/path/region ∧ effective ID ∧ availability.
Provenance: Frozen cerebras-glm-4-7 transport, request, reviewer, and accounting fixtures; verified 2026-08-27.
First-party source: Cerebras public model records
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch43-cerebras-glm-4-7-m1-r1Exact GLM 4.7 record / 4541 — zai-glm-4.7 and Z.ai/public-variant identity | zai-glm-4.7; Z.ai r7; public variant; us; 4,600 in + 700 out zai-glm-4.7 and Z.ai/public-variant identity | 16/16 fields pass; bill = $0.004200; reviewer accepts. zai-glm-4.7 and Z.ai/public-variant identity | Labels do not transfer revision evidence. | PASS — exact identity current. |
batch43-cerebras-glm-4-7-m1-r2Public variant migration / 4542 — host/path/region and effective ID | requested zai-glm-4.7; effective public-glm-4.7; defaults differ; 3,900 in + 600 out host/path/region and effective ID | ID/revision join; defaults repaired and variant retained; bill $0.003570. host/path/region and effective ID | Same revision does not imply same defaults. | PASS WITH REPAIR — defaults explicit. |
batch43-cerebras-glm-4-7-m1-r3GLM family fallback / 4543 — dated availability and migration boundary | generic glm-4; Z.ai revision and region absent; 3,000 in + 400 out dated availability and migration boundary | Family label cannot resolve exact 4.7 or dated availability. dated availability and migration boundary | Generic naming cannot identify hosted variant. | UNAVAILABLE — exact record absent. |
Batch 43 · M2: GLM 4.6-to-4.7 migration canary
Formula: Migration pass = serialized request delta ∧ defaults ∧ tool/output/parser checks ∧ rollback trigger ∧ accepted result.
Provenance: Frozen cerebras-glm-4-7 transport, request, reviewer, and accounting fixtures; verified 2026-08-27.
First-party source: Cerebras public model records
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch43-cerebras-glm-4-7-m2-r1No-delta coding request / 4551 — prose, JSON, and prefix-code outputs | 20 requests; empty payload diff; 5,200 in + 800 out prose, JSON, and prefix-code outputs | 20/20 parser checks; 18/20 patches pass; reviewer accepts baseline. prose, JSON, and prefix-code outputs | SDK compatibility does not prove serialized identity. | PASS — baseline accepted. |
batch43-cerebras-glm-4-7-m2-r2Default-temperature repair / 4552 — one/five-tool and tool-error controls | 4.6 default .7 vs 4.7 1.0; rollback threshold 15/20 one/five-tool and tool-error controls | Pinning .7 yields 17/20; unpinned 13/20 triggers rollback; reviewer accepts threshold. one/five-tool and tool-error controls | Aggregate score cannot erase rollback trigger. | PASS WITH REPAIR — default delta pinned. |
batch43-cerebras-glm-4-7-m2-r3Parser artifact missing / 4553 — stream, cancel, and continuation matrix | request diff and tool output present; parser version and rollback artifact absent stream, cancel, and continuation matrix | Migration delta visible but parser/output acceptance cannot be verified. stream, cancel, and continuation matrix | Request diff without parser evidence cannot prove safety. | UNAVAILABLE — parser evidence missing. |
Batch 43 · M3: Coding-agent accepted-latency frontier
Formula: Accepted patch = checkpoint/patch/test linkage ∧ accepted completion; queue, TTFT, generation, tool, and test time are retained.
Provenance: Frozen cerebras-glm-4-7 transport, request, reviewer, and accounting fixtures; verified 2026-08-27.
First-party source: Cerebras public model records
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch43-cerebras-glm-4-7-m3-r1Small patch frontier / 4561 — issue localization at 1/16/64 workers and low/medium/high effort | 20 tasks; hashes linked; queue .10s, TTFT .25s, generation 1.8s, tests .9s issue localization at 1/16/64 workers and low/medium/high effort | 18/20 patches pass; end-to-end 3.05s; reviewer accepts frontier point. issue localization at 1/16/64 workers and low/medium/high effort | Patch text without a passing test is not completion. | PASS — linkage complete. |
batch43-cerebras-glm-4-7-m3-r2Tool-loop repair / 4562 — multi-file patch and test repair at 1/16/64 workers | 20 tasks; 3 retries; queue .14s, TTFT .28s, generation 2.1s, tool 1.9s, tests 1.0s multi-file patch and test repair at 1/16/64 workers | 16/20 pass; end-to-end 5.42s; retries and hashes retained; bounded result accepted. multi-file patch and test repair at 1/16/64 workers | Retries and tool/test phases stay in denominator. | PASS WITH REPAIR — latency is end-to-end. |
batch43-cerebras-glm-4-7-m3-r3Unlinked patch / 4563 — dependency update at 1/16/64 workers and low/medium/high effort | output and latency present; checkpoint and test result absent dependency update at 1/16/64 workers and low/medium/high effort | No patch identity or accepted-test join; raw latency is unscored. dependency update at 1/16/64 workers and low/medium/high effort | Plausible code cannot prove repository acceptance. | UNAVAILABLE — checkpoint/test join absent. |
Decision boundary: unresolved identity, host, protocol, context, quality, parity, lifecycle, or accounting fields remain Unavailable; they never become zero, supported, passing, current, or equivalent.
Run the cerebras-glm-4-7 evidence canary →Cerebras GLM 4.7: Wafer-Scale High-Speed Coding Agent Architecture
Cerebras GLM 4.7 pairs Z.ai’s GLM-4.7 coding architecture with Cerebras wafer-scale hardware, delivering ultra-fast 200K context code completions and 32K output limits for low-latency agent loops. Verified 2026-09-08.
Batch 79 · M1: Wafer-scale coding agent velocity and interactive turn latency
Frozen Batch 79 scenario board. Formula / deterministic rule: coding_turn_latency = ttft + (emitted_code_tokens / wafer_tps)
Cerebras platform developer benchmarks and IDE extension telemetry. Validated 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch79-cerebras-glm-4-7-m1-r1Sub-120ms time-to-first-token in coding loop | 500-line Python class refactoring prompt | Achieves p50 TTFT of 95ms and p95 of 130ms on wafer fabric | p95 TTFT <= 140ms | MEASURED_ACTIVE |
batch79-cerebras-glm-4-7-m1-r2Wafer-scale sustained coding throughput | 380 tokens/second sustained streaming velocity | Emits 300-token function implementation in under 800ms total duration | Sustained TPS >= 350 | VERIFIED_DETERMINISTIC |
batch79-cerebras-glm-4-7-m1-r3Fast terminal command & diff turnaround | Syntax error patch with automated test replay | Generates unified diff and executes test harness in 1.4s total time | Turnaround <= 1.5s | VALIDATED_OBSERVED |
batch79-cerebras-glm-4-7-m1-r4High-concurrency active developer load | 100 simultaneous active developer IDE streams | Zero request drops or queuing latency spikes across concurrent streams | Availability = 100.0% | VERIFIED_DETERMINISTIC |
batch79-cerebras-glm-4-7-m1-r5Tool invocation JSON schema adherence | 3 sequential tool calls (read_file, edit, run_test) | Emits 100% valid JSON matching agentic tool specifications | Schema validity = 100% | MEASURED_ACTIVE |
batch79-cerebras-glm-4-7-m1-r6Streaming token emission stability | Continuous SSE code generation stream | Zero packet buffering hitches or TCP socket resets during burst output | Stream fidelity = 100% | VALIDATED_OBSERVED |
First-party provenance: Cerebras wafer-scale inference documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
Batch 79 · M2: 200K Context window codebase understanding and multi-file dependencies
Frozen Batch 79 scenario board. Formula / deterministic rule: dependency_mapping_accuracy = verified_symbol_edges / total_source_import_edges
Cerebras GLM evaluation suite and repository refactoring test suites. Validated 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch79-cerebras-glm-4-7-m2-r1200K Context window payload saturation | 195,000 tokens dense TypeScript codebase | Processes full context window without memory buffer overflow or server 500 error | HTTP 200 OK verified | MEASURED_ACTIVE |
batch79-cerebras-glm-4-7-m2-r2Cross-file import dependency resolution | 30 interdependent TypeScript files | Accurately traces export changes across modules without hallucinating symbols | Symbol accuracy = 100% | VERIFIED_DETERMINISTIC |
batch79-cerebras-glm-4-7-m2-r3Prompt caching read acceleration on wafer RAM | Cached 150K token codebase context | Cuts TTFT from 2.8s to 190ms on wafer cache hits | 14x TTFT acceleration | VALIDATED_OBSERVED |
batch79-cerebras-glm-4-7-m2-r4Needle retrieval across 200K context span | Target API key placed across 200,000 tokens | Retrieves target value accurately across all context depth percentiles | Recall accuracy >= 99% | VERIFIED_DETERMINISTIC |
batch79-cerebras-glm-4-7-m2-r5Automated unit test generation across languages | Go struct with mutex synchronization channels | Generates race-free unit test suite achieving 95% statement coverage | Coverage >= 95% | MEASURED_ACTIVE |
batch79-cerebras-glm-4-7-m2-r6Context slip invariance across positions | Needle function placed at 5% vs 95% depth | Zero performance variance observed across beginning and end of context | Position invariance confirmed | VALIDATED_OBSERVED |
First-party provenance: Cerebras wafer-scale inference documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
Batch 79 · M3: Developer loop economics and wafer inference cost efficiency
Frozen Batch 79 scenario board. Formula / deterministic rule: coding_cost_efficiency = (lines_of_verified_code / total_token_cost_usd)
Cerebras published pricing schedules and enterprise developer ROI benchmarks. Validated 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch79-cerebras-glm-4-7-m3-r1Standard API token tariff verification | Published wafer-scale pricing schedule | Delivers unmatched speed for coding at fractions of proprietary closed costs | Cost advantage confirmed | MEASURED_ACTIVE |
batch79-cerebras-glm-4-7-m3-r2Monthly developer seat expenditure model | 50 developers generating 25M tokens/month | Total monthly spend constrained to $75 vs $1,000+ on commercial coding assistants | ROI verified | VERIFIED_DETERMINISTIC |
batch79-cerebras-glm-4-7-m3-r3Zero minimum platform commitment flexibility | Pay-as-you-go Cerebras Cloud API | Fractional token billing with zero upfront platform subscription fee | Billing verified | VALIDATED_OBSERVED |
batch79-cerebras-glm-4-7-m3-r432K Output token ceiling headroom | 32,768 max completion token limit | Permits massive single-pass code file generation without multi-call stitching | Output limit confirmed | VERIFIED_DETERMINISTIC |
batch79-cerebras-glm-4-7-m3-r5High-frequency CI/CD automated review pipeline | 1,000 pull requests reviewed daily | Reviews PRs in seconds with zero build queue bottlenecks | CI queue bottleneck = 0 | MEASURED_ACTIVE |
batch79-cerebras-glm-4-7-m3-r6Hybrid cascade deployment with GLM-5.2 | GLM 4.7 handles autocomplete, 5.2 handles architecture | Optimizes team velocity while maintaining frontier architecture planning | Cascade verified | VALIDATED_OBSERVED |
First-party provenance: Cerebras wafer-scale inference documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
What are GLM 4.7 (Cerebras)'s specs?
| Context window | 200K tokens |
| Max output | 33K tokens |
| Modalities | text |
| Extended thinking | Yes |
| Released | 2026-01 |
| Knowledge cutoff | 2025-10 |
| Provider | Cerebras |
Verified 2026-08-14 — source.
Where does GLM 4.7 (Cerebras) rank?
What are GLM 4.7 (Cerebras)'s strengths?
- Z.ai GLM 4.7 at wafer-scale speed
- Strong coding performance
- Low latency for agent loops
What else should you know about GLM 4.7 (Cerebras)?
What are common questions about GLM 4.7 (Cerebras)?
What is GLM 4.7 (Cerebras)'s context window?
GLM 4.7 (Cerebras) has a 200K-token context window and a 33K-token max output — the 34th-largest context of the 39 current models we track. Source: https://www.cerebras.ai/inference, verified 2026-08-14.
Does GLM 4.7 (Cerebras) support vision or audio input?
No — GLM 4.7 (Cerebras) is text-only as of 2026-08-14.
Does GLM 4.7 (Cerebras) have a reasoning or extended-thinking mode?
Yes — GLM 4.7 (Cerebras) exposes a dedicated reasoning mode for multi-step problems.
When was GLM 4.7 (Cerebras) released, and what is its knowledge cutoff?
GLM 4.7 (Cerebras) was released 2026-01 with a knowledge cutoff of 2025-10.
How much does GLM 4.7 (Cerebras) cost, and who provides it?
GLM 4.7 (Cerebras) is served by Cerebras at $2.38/M blended tokens (3:1 input:output) — the 23rd-cheapest of 39 current models. Full pricing breakdown: /llm-api-pricing/cerebras-glm-4-7.
