← Back to blog
AiAbout 7 min read

Zhipu's GLM-5.2 Goes Fully Open: 744 Billion Parameters, Trained on Domestic Chips

Published Oct 7, 2026
Zhipu's GLM-5.2 Goes Fully Open: 744 Billion Parameters, Trained on Domestic Chips

Zhipu AI, the Beijing lab also known as Z.AI, launched GLM-5.2 and shipped the full open weights within days. The model carries 744 billion total parameters with roughly 40 billion active per token, using a mixture-of-experts architecture with a distinctive sparse-attention optimization called IndexShare.

The licence is MIT. That is the headline for anyone planning to build on it: no territory carve-outs, no revenue trigger, no monthly-active-user cap. Commercial deployment, modification, and redistribution require no additional permission.

What IndexShare buys you

The IndexShare architecture reuses sparse-attention indexers across layers, which the technical write-ups put at a 2.9x reduction in per-token compute at the model's full one-million-token context window. That number matters less as an isolated benchmark and more as an economic fact. Long-context agent sessions are where token costs accumulate, and a 2.9x reduction at the top of the context window changes which workloads are viable on a given cluster.

GLM-5.2 natively supports one million tokens and registers as the top open-weight model on the Artificial Analysis Intelligence Index. On SWE-bench Pro it scored 62.1 percent, ahead of GPT-5.5's 58.6 percent. On Terminal-Bench 2.1 it posted 81.0, the strongest reported score of any open-weight model.

The hardware story is the geopolitical one

The part of GLM-5.2 that will outlast the benchmark chatter is where it runs. The entire GLM-5 series was trained on Huawei Ascend chips using the MindSpore framework. That is a claim no Western lab can make, and it is the reason the model draws attention from buyers who think about supply-chain resilience rather than raw scores.

Zhipu's release materials add that online inference runs across multiple domestic compute platforms, with day-zero compatibility for Huawei Ascend, T-Head, Moore Threads, Cambricon, Kunlun, Muxi, Hygon, and Biren. The company frames this as stable operation at high throughput, low latency, and large concurrency on domestic chip clusters.

For an enterprise outside China, the practical question is whether the weights give a self-hosted path that removes per-token costs. For an enterprise inside China, the question is broader: whether a frontier-capable model can be built and served without depending on hardware it cannot reliably buy.

What it costs, and what self-hosting changes

The pricing tiers separate cleanly. Direct API access runs about $1.40 per million input tokens and $4.40 per million output tokens, with cached input at $0.26. Through OpenRouter, the rates are roughly $1.00 and $4.00. The GLM Coding Plan subscription sits at $10 to $18 per month for a developer-focused tier, which is not meant for production scale.

Self-hosting is where the arithmetic turns. On-premise deployment carries infrastructure cost and no per-token fees. For a workload processing millions of tokens a day, that difference is substantial. The tradeoff is the GPU cluster, the maintenance, and the scaling work that a managed API hides. The MIT licence reduces the legal friction on top of the economic one.

An open geometric wireframe cube of thin brass rods resting on a pale concrete plinth

Quantization is the other lever. With quantized weights, the model reportedly runs on consumer-grade hardware, which opens a self-hosted path for privacy-sensitive applications such as agents handling medical records or proprietary codebases.

The security footnote worth reading

One finding deserves attention before anyone treats open weights as a free lunch. Anthropic's red-team evaluation reported that GLM-5.3 exhibits Mythos-class autonomous cyber capabilities with weak safeguards against open-weight misuse. Open weights cannot be recalled. Once downloaded, the model stays in the deploying organization's possession indefinitely.

That creates a specific deployment problem. An open-weight model running inside an agent framework needs containment layers that a closed API provider handles on its own side. The MIT licence removes vendor lock-in, and it also removes vendor-provided security guarantees. Teams deploying GLM-5.2 in an agent loop should budget for runtime containment, not treat it as someone else's problem.

Why this release lands differently

Plenty of open models ship each month. GLM-5.2 lands differently for three reasons that compound.

The licence is genuinely permissive at a parameter scale where that is rare. Most frontier-class open releases come with a commercial condition, a territory limit, or a user cap. GLM-5.2 has none of those.

The hardware path is independent of NVIDIA. For buyers who treat supply-chain risk as a first-order concern, that changes the shortlist regardless of benchmark rank.

And the compute efficiency at long context makes sustained multi-step sessions economically viable on smaller clusters than a model like NVIDIA's Nemotron 3 Ultra, which needs 8xH100 nodes at its documented minimum.

GLM-5.2 is not the best model at every task, and the security caveat is real. But as a signal of what an open release can now be, frontier-competitive, permissively licensed, and built without US-manufactured GPUs, it is the clearest one this cycle.

The family around the flagship

It helps to see GLM-5.2 as the top of a ladder rather than a standalone release. The GLM family spans from the lightweight GLM-4.6V-Flash at 9B, aimed at edge deployment, up to the 744B flagship. Specialized variants include GLM-5V-Turbo for multimodal vision tasks and the AutoGLM agent framework for complex multi-step planning.

The permissive licence applies across that ladder, which is the part that matters for teams choosing a stack. A company can prototype on the small edge model, scale to the flagship, and move between them without renegotiating rights at each step. That continuity is unusual in the open ecosystem, where licences often differ between a lab's small and large releases.

What the efficiency claim means in practice

The 2.9x reduction in per-token compute at full context is the kind of number that reads as a spec until you attach it to a workload.

Consider an agent that reasons over a large codebase or a long document set. Each step re-reads a growing context, and the cost of that re-reading is what makes long sessions expensive. A reduction of nearly threefold at the top of the context window lowers the point at which a long-running session stops being worth it. The session still costs money, but the break-even moves.

That is also why the quantization path matters. With quantized weights, the model reportedly runs on consumer-grade hardware, which turns a long-context agent from something that needs a cluster into something a small team can run locally. For an application handling sensitive data, such as medical records or proprietary source code, running locally is the only deployment that fits the compliance constraint, and cost is a secondary benefit.

Where the security caveat bites

The red-team finding is easy to file away as someone else's problem. It is not, and the reason is structural.

A closed API provider can update a model, change a policy, or cut off a misuse case after the fact. An open-weight release removes that lever. The weights, once downloaded, stay in the deploying organization's possession, and no vendor can revoke them. That is the whole point of open weights, and it is also why containment has to be built by the team that deploys them.

In practice that means an agent running GLM-5.2 needs the same runtime discipline a team would apply to any powerful tool with write access: scoped permissions, a sandbox for untrusted actions, logging that survives a crash, and a human escalation path for consequential decisions. None of that is provided by the licence, and none of it comes free with the weights. The MIT terms remove vendor lock-in on the legal side and, by design, remove vendor-provided safety on the operational side.

Related articles