GLM-5.2
Long-horizon coding on open weights at a fraction of frontier pricing.
GLM-5.2 supersedes GLM-5.1.
What are GLM-5.2's specs and price?
GLM-5.2, built by Z.ai, ships a 1M-token context window and a 64K-token max output, released 2026-05. It supports text input with a dedicated reasoning mode and costs $2.15 per million blended tokens, the 21st-cheapest of 39 models we track.
Batch 44 evidence surface · verified 2026-08-14 · frozen route allowlist: /models/glm-5-2
GLM-5.2 exact identity, request envelope, and runtime evidence
Batch 44 · M1: GLM-5.2 artifact-and-surface identity ledger
Formula / rubric: artifact pass = BF16/W8A8/W4A8C8 artifact ∧ hosted surface ∧ revision join.
Dated provenance: Frozen Batch 44 models-glm-5-2 fixture; GLM-5.2 hosted artifact chain; reviewer ledger verified 2026-08-14.
First-party citation: Z.ai GLM-5.2 model guide
| Field ID / fixture | Inputs | Observation / calculation | Decision boundary | State |
|---|---|---|---|---|
batch44-models-glm-5-2-m1-r1BF16 hosted artifact chain | BF16 artifact; repository commit; hosted endpoint; revision; request ID | BF16 artifact, commit, and hosted endpoint form one attributable chain. | A repository tag alone does not prove hosted weights. | PASS — chain joined. |
batch44-models-glm-5-2-m1-r2W8A8 hosted artifact chain | W8A8 artifact; quantization manifest; host; revision; tokenizer | W8A8 is retained as a distinct hosted artifact with its own manifest. | Quantized artifacts cannot inherit BF16 benchmark evidence. | PASS — chain joined. |
batch44-models-glm-5-2-m1-r3W4A8C8 hosted artifact chain | W4A8C8 artifact; calibration file; host; revision; checksum | The hosted artifact is named, but the calibration checksum is missing. | Missing calibration identity blocks artifact attribution. | UNAVAILABLE — chain incomplete. |
Batch 44 · M2: Reasoning-control and tool-order matrix
Formula / rubric: order pass = reasoning control state ∧ tool-call order ∧ result settlement.
Dated provenance: Frozen Batch 44 models-glm-5-2 fixture; GLM-5.2 reasoning and tool-order fixtures; reviewer ledger verified 2026-08-14.
First-party citation: Z.ai GLM-5.2 model guide
| Field ID / fixture | Inputs | Observation / calculation | Decision boundary | State |
|---|---|---|---|---|
batch44-models-glm-5-2-m2-r1reasoning then tool | reasoning enabled; tool call 1; tool result 1; final answer; event order | Reasoning control precedes the tool call and the ordered result settles. | Final text quality cannot repair event-order loss. | PASS — order preserved. |
batch44-models-glm-5-2-m2-r2tool then reasoning retry | tool call; retry; reasoning control; call IDs; usage; cancellation | Retry is recorded after the tool result; reasoning order is explicit. | Retry order must remain attributable to the original tool call. | PASS WITH REPAIR — retry linked. |
batch44-models-glm-5-2-m2-r3reasoning/tool conflict | reasoning flag; parallel tools; returned events; missing stop reason | Tool results arrive, but the missing stop reason prevents settlement. | Tool completion without stop state is not accepted. | UNAVAILABLE — stop join missing. |
Batch 44 · M3: Long-horizon coding checkpoint ledger
Formula / rubric: checkpoint = 100K/500K/850K/near-1M inventory item with artifact, tool, test, and billing joins.
Dated provenance: Frozen Batch 44 models-glm-5-2 fixture; GLM-5.2 long-horizon checkpoint inventory; reviewer ledger verified 2026-08-14.
First-party citation: vLLM Ascend GLM-5.2 deployment guide
| Field ID / fixture | Inputs | Observation / calculation | Decision boundary | State |
|---|---|---|---|---|
batch44-models-glm-5-2-m3-r1100K checkpoint | 100K repository tokens; checkpoint C100; tool calls; tests; usage | Checkpoint C100 restores and all tests settle with a billing join. | Short checkpoint success does not establish long-horizon continuity. | PASS — checkpoint settled. |
batch44-models-glm-5-2-m3-r2500K/850K checkpoints | 500K checkpoint C500; 850K checkpoint C850; tool order; test artifacts | C500 passes; C850 has a missing test artifact and remains unavailable. | Checkpoint inventory is per size, not interpolated. | UNAVAILABLE — C850 artifact missing. |
batch44-models-glm-5-2-m3-r3near-1M checkpoint | near-1M repository; resume event; BF16/W8A8/W4A8C8 IDs; rollback owner | Near-1M resume reaches the boundary, but quantization-specific billing is not joined. | Near-limit completion cannot inherit artifact-chain or cost evidence. | UNAVAILABLE — billing chain missing. |
Fail-closed rule: an unresolved identity, host, endpoint, artifact, modality, control, workload, acceptance, or accounting join remains Unavailable; no fallback or neighboring route supplies it.
Run the models-glm-5-2 evidence canary →GLM 5.2: Z.ai Open-Weights Coding Flagship & 1M Context Architecture
GLM 5.2 is Z.ai’s coding-first open-weights flagship under an MIT license, featuring 1,000,000 token context window, 64K max output, and deep reasoning that beats larger frontier models on long-horizon coding. Verified 2026-09-08.
Batch 79 · M1: Long-horizon software engineering and multi-file repository refactoring
Frozen Batch 79 scenario board. Formula / deterministic rule: swe_pass_rate = passing_github_benchmarks / total_evaluated_real_world_issues
Z.ai SWE-bench evaluations and open-weights benchmark test suites. Validated 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch79-glm-5-2-m1-r1SWE-bench verified real-world issue resolution | Complex Django / CPython GitHub issues | Resolves difficult issues beating closed frontier models on multiple coding tasks | Pass rate verified | MEASURED_ACTIVE |
batch79-glm-5-2-m1-r2Full-stack repository architecture migration | Migrating REST API to tRPC end-to-end types | Rewrites 35 router files and updates client hooks with 100% type safety | Zero typecheck errors | VERIFIED_DETERMINISTIC |
batch79-glm-5-2-m1-r3Reasoning mode code synthesis verification | NP-hard scheduling algorithm with constraint solver | Utilizes extended thinking tokens and outputs proven optimal heuristic | Solution provably valid | VALIDATED_OBSERVED |
batch79-glm-5-2-m1-r4Autonomous unit test regression generation | Edge case branch coverage for payment gateway | Generates 50 parameterized tests achieving 98% branch coverage | Coverage >= 95% | VERIFIED_DETERMINISTIC |
batch79-glm-5-2-m1-r5MIT license open weights audit compliance | Hugging Face weights inspection | Completely open for commercial modification, self-hosting, and private deployment | MIT license verified | MEASURED_ACTIVE |
batch79-glm-5-2-m1-r6Streaming code completion velocity | 75 tokens/second steady generation rate | Smooth text emission across continuous coding sessions | Steady TPS >= 70 | VALIDATED_OBSERVED |
First-party provenance: Z.ai GLM platform documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
Batch 79 · M2: 1M Context window processing and deep codebase navigation
Frozen Batch 79 scenario board. Formula / deterministic rule: codebase_indexing_f1 = (2 · precision · recall) / (precision + recall)
Z.ai 1M context evaluation suite and enterprise repository benchmarks. Validated 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch79-glm-5-2-m2-r11M Context full payload capacity saturation | 1,000,000 tokens dense repository payload | Processes full context window without memory exhaustion or server 500 error | HTTP 200 OK verified | MEASURED_ACTIVE |
batch79-glm-5-2-m2-r2Needle retrieval across 1M context span | 50 distinct function signatures across 1M tokens | Retrieves 50/50 signatures with exact line and file citations | Recall = 100.0% | VERIFIED_DETERMINISTIC |
batch79-glm-5-2-m2-r3Prompt caching 75% discount on 1M context | Cached 800K token codebase context | Reduces input token price from $1.40/M to $0.35/M on prompt cache hits | 75% discount applied | VALIDATED_OBSERVED |
batch79-glm-5-2-m2-r4Cross-module symbol refactoring across 1M context | Renaming core database entity across 100 files | Updates all 350 call sites without missing a single reference | Refactor completeness = 100% | VERIFIED_DETERMINISTIC |
batch79-glm-5-2-m2-r5Prompt cache TTFT acceleration | Cached 600K token repository preamble | Cuts TTFT from 18s to 1.2s on prompt cache hit | 15x TTFT acceleration | MEASURED_ACTIVE |
batch79-glm-5-2-m2-r6Context slip invariance across positions | Target key positioned at 1%, 50%, and 99% depth | Zero variance in recall accuracy across token depth percentiles | Position invariance confirmed | VALIDATED_OBSERVED |
First-party provenance: Z.ai GLM platform documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
Batch 79 · M3: Open weights economics: cloud API vs on-premises GPU datacenter TCO
Frozen Batch 79 scenario board. Formula / deterministic rule: tco_comparison = (cloud_api_annual_spend - datacenter_annual_amortization)
Z.ai published API pricing schedules and enterprise hardware TCO models. Validated 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch79-glm-5-2-m3-r1Standard API token tariff verification | $1.40/M input, $4.40/M output tariffs | Delivers frontier coding intelligence at 70% lower price than closed flagships | Cost advantage confirmed | MEASURED_ACTIVE |
batch79-glm-5-2-m3-r2High-volume production spend modeling | 2 billion tokens monthly throughput | Total spend under $4,500 vs $15,000+ on commercial proprietary flagships | ROI confirmed | VERIFIED_DETERMINISTIC |
batch79-glm-5-2-m3-r3Private air-gapped on-premises deployment | Self-hosted on 4x H100 GPU nodes | Guarantees proprietary source code never leaves private enterprise perimeter | Data sovereignty = 100% | VALIDATED_OBSERVED |
batch79-glm-5-2-m3-r464K Output token ceiling headroom | 64,000 max completion token limit | Generates entire software modules in single continuous generation pass | Output limit confirmed | VERIFIED_DETERMINISTIC |
batch79-glm-5-2-m3-r5Zero enterprise lock-in elasticity | MIT license weights portability | Freedom to switch between cloud API, vLLM, TensorRT-LLM, and on-prem hardware | Portability confirmed | MEASURED_ACTIVE |
batch79-glm-5-2-m3-r6Self-hosted cluster break-even threshold | Cloud API vs owned 4x H100 server | Cloud API remains more economical than owned hardware below 120M tok/day | Break-even confirmed | VALIDATED_OBSERVED |
First-party provenance: Z.ai GLM platform documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
What are GLM-5.2's specs?
| Context window | 1M tokens |
| Max output | 64K tokens |
| Modalities | text |
| Extended thinking | Yes |
| Released | 2026-05 |
| Knowledge cutoff | 2026-02 |
| Provider | Z.ai |
Verified 2026-08-14 — source.
Where does GLM-5.2 rank?
What are GLM-5.2's strengths?
- Open-weights coding-first flagship
- 1M-token context window
- Beats larger frontier models on long-horizon coding
What else should you know about GLM-5.2?
What are common questions about GLM-5.2?
What is GLM-5.2's context window?
GLM-5.2 has a 1M-token context window and a 64K-token max output — the 17th-largest context of the 39 current models we track. Source: https://docs.z.ai/, verified 2026-08-14.
Does GLM-5.2 support vision or audio input?
No — GLM-5.2 is text-only as of 2026-08-14.
Does GLM-5.2 have a reasoning or extended-thinking mode?
Yes — GLM-5.2 exposes a dedicated reasoning mode for multi-step problems.
When was GLM-5.2 released, and what is its knowledge cutoff?
GLM-5.2 was released 2026-05 with a knowledge cutoff of 2026-02.
How much does GLM-5.2 cost, and who provides it?
GLM-5.2 is served by Z.ai at $2.15/M blended tokens (3:1 input:output) — the 21st-cheapest of 39 current models. Full pricing breakdown: /llm-api-pricing/glm-5-2.
