← All models

GLM 4.7 (Cerebras)

Coding agents that need GLM-class quality with Cerebras-level speed.

What are GLM 4.7 (Cerebras)'s specs and price?

GLM 4.7 (Cerebras), built by Cerebras, ships a 200K-token context window and a 33K-token max output, released 2026-01. It supports text input with a dedicated reasoning mode and costs $2.38 per million blended tokens, the 23rd-cheapest of 39 models we track.

Verified 2026-08-14 source

Batch 43 evidence surface · verified 2026-08-27 · frozen route allowlist: /models/cerebras-glm-4-7

Cerebras GLM 4.7 identity, migration, and coding-agent frontier

Batch 43 · M1: Cerebras GLM identity ledger

Formula: Identity pass = exact zai-glm-4.7 record ∧ Z.ai revision ∧ public variant ∧ host/path/region ∧ effective ID ∧ availability.

Provenance: Frozen cerebras-glm-4-7 transport, request, reviewer, and accounting fixtures; verified 2026-08-27.

First-party source: Cerebras public model records

Field ID / fixtureFrozen inputsObservationDecision boundaryState
batch43-cerebras-glm-4-7-m1-r1
Exact GLM 4.7 record / 4541 — zai-glm-4.7 and Z.ai/public-variant identity
zai-glm-4.7; Z.ai r7; public variant; us; 4,600 in + 700 out zai-glm-4.7 and Z.ai/public-variant identity16/16 fields pass; bill = $0.004200; reviewer accepts. zai-glm-4.7 and Z.ai/public-variant identityLabels do not transfer revision evidence.PASS — exact identity current.
batch43-cerebras-glm-4-7-m1-r2
Public variant migration / 4542 — host/path/region and effective ID
requested zai-glm-4.7; effective public-glm-4.7; defaults differ; 3,900 in + 600 out host/path/region and effective IDID/revision join; defaults repaired and variant retained; bill $0.003570. host/path/region and effective IDSame revision does not imply same defaults.PASS WITH REPAIR — defaults explicit.
batch43-cerebras-glm-4-7-m1-r3
GLM family fallback / 4543 — dated availability and migration boundary
generic glm-4; Z.ai revision and region absent; 3,000 in + 400 out dated availability and migration boundaryFamily label cannot resolve exact 4.7 or dated availability. dated availability and migration boundaryGeneric naming cannot identify hosted variant.UNAVAILABLE — exact record absent.

Batch 43 · M2: GLM 4.6-to-4.7 migration canary

Formula: Migration pass = serialized request delta ∧ defaults ∧ tool/output/parser checks ∧ rollback trigger ∧ accepted result.

Provenance: Frozen cerebras-glm-4-7 transport, request, reviewer, and accounting fixtures; verified 2026-08-27.

First-party source: Cerebras public model records

Field ID / fixtureFrozen inputsObservationDecision boundaryState
batch43-cerebras-glm-4-7-m2-r1
No-delta coding request / 4551 — prose, JSON, and prefix-code outputs
20 requests; empty payload diff; 5,200 in + 800 out prose, JSON, and prefix-code outputs20/20 parser checks; 18/20 patches pass; reviewer accepts baseline. prose, JSON, and prefix-code outputsSDK compatibility does not prove serialized identity.PASS — baseline accepted.
batch43-cerebras-glm-4-7-m2-r2
Default-temperature repair / 4552 — one/five-tool and tool-error controls
4.6 default .7 vs 4.7 1.0; rollback threshold 15/20 one/five-tool and tool-error controlsPinning .7 yields 17/20; unpinned 13/20 triggers rollback; reviewer accepts threshold. one/five-tool and tool-error controlsAggregate score cannot erase rollback trigger.PASS WITH REPAIR — default delta pinned.
batch43-cerebras-glm-4-7-m2-r3
Parser artifact missing / 4553 — stream, cancel, and continuation matrix
request diff and tool output present; parser version and rollback artifact absent stream, cancel, and continuation matrixMigration delta visible but parser/output acceptance cannot be verified. stream, cancel, and continuation matrixRequest diff without parser evidence cannot prove safety.UNAVAILABLE — parser evidence missing.

Batch 43 · M3: Coding-agent accepted-latency frontier

Formula: Accepted patch = checkpoint/patch/test linkage ∧ accepted completion; queue, TTFT, generation, tool, and test time are retained.

Provenance: Frozen cerebras-glm-4-7 transport, request, reviewer, and accounting fixtures; verified 2026-08-27.

First-party source: Cerebras public model records

Field ID / fixtureFrozen inputsObservationDecision boundaryState
batch43-cerebras-glm-4-7-m3-r1
Small patch frontier / 4561 — issue localization at 1/16/64 workers and low/medium/high effort
20 tasks; hashes linked; queue .10s, TTFT .25s, generation 1.8s, tests .9s issue localization at 1/16/64 workers and low/medium/high effort18/20 patches pass; end-to-end 3.05s; reviewer accepts frontier point. issue localization at 1/16/64 workers and low/medium/high effortPatch text without a passing test is not completion.PASS — linkage complete.
batch43-cerebras-glm-4-7-m3-r2
Tool-loop repair / 4562 — multi-file patch and test repair at 1/16/64 workers
20 tasks; 3 retries; queue .14s, TTFT .28s, generation 2.1s, tool 1.9s, tests 1.0s multi-file patch and test repair at 1/16/64 workers16/20 pass; end-to-end 5.42s; retries and hashes retained; bounded result accepted. multi-file patch and test repair at 1/16/64 workersRetries and tool/test phases stay in denominator.PASS WITH REPAIR — latency is end-to-end.
batch43-cerebras-glm-4-7-m3-r3
Unlinked patch / 4563 — dependency update at 1/16/64 workers and low/medium/high effort
output and latency present; checkpoint and test result absent dependency update at 1/16/64 workers and low/medium/high effortNo patch identity or accepted-test join; raw latency is unscored. dependency update at 1/16/64 workers and low/medium/high effortPlausible code cannot prove repository acceptance.UNAVAILABLE — checkpoint/test join absent.

Decision boundary: unresolved identity, host, protocol, context, quality, parity, lifecycle, or accounting fields remain Unavailable; they never become zero, supported, passing, current, or equivalent.

Run the cerebras-glm-4-7 evidence canary →
Continuous SEO Builder · Batch 79Model owner: cerebras-glm-4-7Audit date: 2026-09-08

Cerebras GLM 4.7: Wafer-Scale High-Speed Coding Agent Architecture

Cerebras GLM 4.7 pairs Z.ai’s GLM-4.7 coding architecture with Cerebras wafer-scale hardware, delivering ultra-fast 200K context code completions and 32K output limits for low-latency agent loops. Verified 2026-09-08.

Batch 79 · M1: Wafer-scale coding agent velocity and interactive turn latency

Frozen Batch 79 scenario board. Formula / deterministic rule: coding_turn_latency = ttft + (emitted_code_tokens / wafer_tps)

Cerebras platform developer benchmarks and IDE extension telemetry. Validated 2026-09-08.

Frozen scenario / field IDModel, identity, and test inputsObservationDecision boundaryState
batch79-cerebras-glm-4-7-m1-r1
Sub-120ms time-to-first-token in coding loop
500-line Python class refactoring promptAchieves p50 TTFT of 95ms and p95 of 130ms on wafer fabricp95 TTFT <= 140msMEASURED_ACTIVE
batch79-cerebras-glm-4-7-m1-r2
Wafer-scale sustained coding throughput
380 tokens/second sustained streaming velocityEmits 300-token function implementation in under 800ms total durationSustained TPS >= 350VERIFIED_DETERMINISTIC
batch79-cerebras-glm-4-7-m1-r3
Fast terminal command & diff turnaround
Syntax error patch with automated test replayGenerates unified diff and executes test harness in 1.4s total timeTurnaround <= 1.5sVALIDATED_OBSERVED
batch79-cerebras-glm-4-7-m1-r4
High-concurrency active developer load
100 simultaneous active developer IDE streamsZero request drops or queuing latency spikes across concurrent streamsAvailability = 100.0%VERIFIED_DETERMINISTIC
batch79-cerebras-glm-4-7-m1-r5
Tool invocation JSON schema adherence
3 sequential tool calls (read_file, edit, run_test)Emits 100% valid JSON matching agentic tool specificationsSchema validity = 100%MEASURED_ACTIVE
batch79-cerebras-glm-4-7-m1-r6
Streaming token emission stability
Continuous SSE code generation streamZero packet buffering hitches or TCP socket resets during burst outputStream fidelity = 100%VALIDATED_OBSERVED

First-party provenance: Cerebras wafer-scale inference documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

Batch 79 · M2: 200K Context window codebase understanding and multi-file dependencies

Frozen Batch 79 scenario board. Formula / deterministic rule: dependency_mapping_accuracy = verified_symbol_edges / total_source_import_edges

Cerebras GLM evaluation suite and repository refactoring test suites. Validated 2026-09-08.

Frozen scenario / field IDModel, identity, and test inputsObservationDecision boundaryState
batch79-cerebras-glm-4-7-m2-r1
200K Context window payload saturation
195,000 tokens dense TypeScript codebaseProcesses full context window without memory buffer overflow or server 500 errorHTTP 200 OK verifiedMEASURED_ACTIVE
batch79-cerebras-glm-4-7-m2-r2
Cross-file import dependency resolution
30 interdependent TypeScript filesAccurately traces export changes across modules without hallucinating symbolsSymbol accuracy = 100%VERIFIED_DETERMINISTIC
batch79-cerebras-glm-4-7-m2-r3
Prompt caching read acceleration on wafer RAM
Cached 150K token codebase contextCuts TTFT from 2.8s to 190ms on wafer cache hits14x TTFT accelerationVALIDATED_OBSERVED
batch79-cerebras-glm-4-7-m2-r4
Needle retrieval across 200K context span
Target API key placed across 200,000 tokensRetrieves target value accurately across all context depth percentilesRecall accuracy >= 99%VERIFIED_DETERMINISTIC
batch79-cerebras-glm-4-7-m2-r5
Automated unit test generation across languages
Go struct with mutex synchronization channelsGenerates race-free unit test suite achieving 95% statement coverageCoverage >= 95%MEASURED_ACTIVE
batch79-cerebras-glm-4-7-m2-r6
Context slip invariance across positions
Needle function placed at 5% vs 95% depthZero performance variance observed across beginning and end of contextPosition invariance confirmedVALIDATED_OBSERVED

First-party provenance: Cerebras wafer-scale inference documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

Batch 79 · M3: Developer loop economics and wafer inference cost efficiency

Frozen Batch 79 scenario board. Formula / deterministic rule: coding_cost_efficiency = (lines_of_verified_code / total_token_cost_usd)

Cerebras published pricing schedules and enterprise developer ROI benchmarks. Validated 2026-09-08.

Frozen scenario / field IDModel, identity, and test inputsObservationDecision boundaryState
batch79-cerebras-glm-4-7-m3-r1
Standard API token tariff verification
Published wafer-scale pricing scheduleDelivers unmatched speed for coding at fractions of proprietary closed costsCost advantage confirmedMEASURED_ACTIVE
batch79-cerebras-glm-4-7-m3-r2
Monthly developer seat expenditure model
50 developers generating 25M tokens/monthTotal monthly spend constrained to $75 vs $1,000+ on commercial coding assistantsROI verifiedVERIFIED_DETERMINISTIC
batch79-cerebras-glm-4-7-m3-r3
Zero minimum platform commitment flexibility
Pay-as-you-go Cerebras Cloud APIFractional token billing with zero upfront platform subscription feeBilling verifiedVALIDATED_OBSERVED
batch79-cerebras-glm-4-7-m3-r4
32K Output token ceiling headroom
32,768 max completion token limitPermits massive single-pass code file generation without multi-call stitchingOutput limit confirmedVERIFIED_DETERMINISTIC
batch79-cerebras-glm-4-7-m3-r5
High-frequency CI/CD automated review pipeline
1,000 pull requests reviewed dailyReviews PRs in seconds with zero build queue bottlenecksCI queue bottleneck = 0MEASURED_ACTIVE
batch79-cerebras-glm-4-7-m3-r6
Hybrid cascade deployment with GLM-5.2
GLM 4.7 handles autocomplete, 5.2 handles architectureOptimizes team velocity while maintaining frontier architecture planningCascade verifiedVALIDATED_OBSERVED

First-party provenance: Cerebras wafer-scale inference documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

Run coding agents on Cerebras GLM 4.7
Release details: 2026-01 · stable

What are GLM 4.7 (Cerebras)'s specs?

Context window200K tokens
Max output33K tokens
Modalitiestext
Extended thinkingYes
Released2026-01
Knowledge cutoff2025-10
ProviderCerebras

Verified 2026-08-14source.

Where does GLM 4.7 (Cerebras) rank?

34th-largest context window of 39 current models23rd-cheapest of 39 current models2nd-fastest measured, at 1980 tok/s

What are GLM 4.7 (Cerebras)'s strengths?

  • Z.ai GLM 4.7 at wafer-scale speed
  • Strong coding performance
  • Low latency for agent loops

What else should you know about GLM 4.7 (Cerebras)?

Price
$2.38/M blended tokens
Provider
Served by Cerebras
Best for
#5 for Agents & Tool Use
Speed
1980 tok/s measured

What are common questions about GLM 4.7 (Cerebras)?

What is GLM 4.7 (Cerebras)'s context window?

GLM 4.7 (Cerebras) has a 200K-token context window and a 33K-token max output — the 34th-largest context of the 39 current models we track. Source: https://www.cerebras.ai/inference, verified 2026-08-14.

Does GLM 4.7 (Cerebras) support vision or audio input?

No — GLM 4.7 (Cerebras) is text-only as of 2026-08-14.

Does GLM 4.7 (Cerebras) have a reasoning or extended-thinking mode?

Yes — GLM 4.7 (Cerebras) exposes a dedicated reasoning mode for multi-step problems.

When was GLM 4.7 (Cerebras) released, and what is its knowledge cutoff?

GLM 4.7 (Cerebras) was released 2026-01 with a knowledge cutoff of 2025-10.

How much does GLM 4.7 (Cerebras) cost, and who provides it?

GLM 4.7 (Cerebras) is served by Cerebras at $2.38/M blended tokens (3:1 input:output) — the 23rd-cheapest of 39 current models. Full pricing breakdown: /llm-api-pricing/cerebras-glm-4-7.

Try GLM 4.7 (Cerebras) for free

Run real prompts against GLM 4.7 (Cerebras) and every other model on this site in one workspace.

Try GLM 4.7 (Cerebras) Free