GLM-5.3: Z.ai holds back its own model over cyber capability
China's Z.ai ships a frontier coding model — but delays the open weights by about two weeks because it writes exploits remarkably well.

Illustration · AI-generated (AI IN LIFE)
At a glance
- Unveiled August 14, 2026; API immediately, weights about two weeks later
- ExploitBench: 54.4% (GLM-5.2: 24.4%), CyberGym: 84.5%
- 105 exploit tasks solved in two hours (predecessor: 29)
- 2,436 reported vulnerabilities across 269 open-source projects, 53 CVEs published
- OpenVuln: free repo-scanning tool on Hugging Face
What happened? Chinese lab Z.ai unveiled GLM-5.3 — available via API and its coding subscription, but without the usual open weights. The reason is unusually candid: in internal tests, the model showed offensive capabilities far beyond its predecessor. The weights are due only after additional safety reviews, around the end of August.
How big is the jump? On ExploitBench, the score more than doubled from 24.4 to 54.4 percent, and on the CyberGym benchmark for vulnerability discovery GLM-5.3 reached 84.5 percent (predecessor: 77.2). In a two-hour window the model solved 105 exploit tasks — GLM-5.2 managed 29. All figures are company-reported; independent verification is pending.
Can the model defend, too? That is Z.ai's core argument. Since GLM-5.2, the company says it has surfaced 2,436 vulnerabilities across 269 open-source projects, including 1,097 rated critical or high; 53 have already been published with CVEs. It also launched OpenVuln, a Hugging Face tool that lets maintainers scan their repositories with GLM models.
How good is it at coding? Z.ai calls GLM-5.3 "one of the strongest open-weight coding models". On Terminal-Bench 3.0 the score jumps from 4.6 to 28.3, on DeepSWE v1.1 from 46.2 to 66.9. On the hardest evaluations, Western frontier models reportedly stay ahead.
Why does it matter? For the first time, a Chinese lab is explicitly delaying an open-weight release over cyber risk — the same trade-off OpenAI and Anthropic have wrestled with for months. Once weights are freely available, their use cannot be recalled: the same model family then serves attackers and defenders alike.
This article was produced with AI assistance and editorially reviewed.
FAQ
Why is Z.ai holding back the weights?
Internal tests showed sharply increased exploit capability. Z.ai wants to finish extra safety reviews with selected partners before release — scheduled for roughly two weeks.
Can I use GLM-5.3 already?
Yes, via the Z.ai API and the GLM Coding Plan. Only the downloadable weights, license, and model card are still missing.
Are the benchmark numbers independently confirmed?
No. CyberGym, ExploitBench, and the coding scores come from Z.ai's own evaluations; independent retests are still pending.


