On the 84.5% CyberGym score, Z.ai's two-week hold, and whether that timeline is actually enough
Z.ai's GLM-5.3 tops the hacking leaderboard. Sitting on the weights was the right call.
Anti-AI
00
Skeptic
01
Neutral
01
Pro (practical)
02
Pro (hyped)
00
← Anti-AI · Pro-AI →
Somewhere in the software running your bank's website, your hospital records system, or the app you ordered lunch from last Tuesday, there are probably security holes. Most get found slowly — by human teams who spend years learning to look for them, or by luck, or (too often) by someone who found them first and didn't tell anyone.
Z.ai's new model finds them differently.
GLM-5.3 launched on August 14, an update to Z.ai's existing 743 billion-parameter model that kept the same underlying architecture but pushed post-training much harder. The result scored 84.5% on CyberGym, the benchmark the security research field uses to measure how well an AI can discover and demonstrate real software vulnerabilities. That puts it first on the current leaderboard, ahead of Anthropic's Mythos 5 at 83.8% and OpenAI's GPT-5.6 Sol at 83.6%. Z.ai claims the model surfaced 2,436 vulnerabilities across 269 open-source projects, 1,097 rated critical or high severity.
Z.ai also said it is not releasing the model weights. Not yet. The company has committed to roughly a two-week hold to run additional safety evaluations before the checkpoint goes public. Expected release: around August 28.
I think the hold is correct.
Source spread
- Axios — Z.ai holds GLM-5.3 release over hacking risks — [builder]. Primary source on the two-week delay. Confirms Z.ai's stated reasoning and the expected weights date.
- VentureBeat — GLM-5.3 and a 'serious vulnerability' in Cursor — [skeptic]. Reports the model found a serious vulnerability in Cursor during internal evaluation; questions the scope of the 2,436 count.
- BenchLM — CyberGym leaderboard August 2026 — [builder]. Independent score source for the 84.5%, 83.8%, 83.6% comparative rankings.
- D-Central — What the 84.5% score hides — [skeptic]. Raises methodological questions about the self-reported vulnerability count and how it compares to independently verified disclosure.
What's real
- GLM-5.3 scored 84.5% on a benchmark that isn't Z.ai's own. CyberGym is independently maintained; the score isn't self-certified.
- The two-week hold is a meaningful signal. Plenty of open-weight releases have shipped with zero safety review. Choosing to delay costs Z.ai commercially, which is usually the most honest signal of intent.
- The API is live now, through Z.ai's GLM Coding Plan starting at $18 per month. Security researchers can evaluate the model's actual capabilities today without waiting for the weights.
- Post-training improvements are reproducible. If GLM-5.3's cyber capability comes entirely from scaled-up post-training on the same base as GLM-5.2, other labs will understand the technique quickly. This is not a one-lab capability.
What deserves a side-eye
- The 2,436 vulnerabilities figure is Z.ai's own claim over projects Z.ai chose. Independent reproduction hasn't happened yet. The number may be real; it's also the kind of claim that benefits from Z.ai choosing favorable target software.
- Two weeks is a short review window for a model scoring this high on exploit capability. OpenAI spent months evaluating Daybreak before deploying GPT-5.6 Cyber. What exactly does two weeks cover, and what does it not?
- "Critical or high severity" is a scoring-methodology-dependent label. The same flaw can be CVSS 9.8 under one taxonomy and 7.1 under another. The 1,097 count is only as meaningful as the rating methodology behind it.
What to do about it
- If you run software your customers depend on: The threat model changed. AI-powered vulnerability scanning is moving from "large security teams have this" to "anyone with $18/month and a browser has this." Bug bounty programs, patch velocity, and automated dependency scanning matter more than they did a month ago.
- If you're a security practitioner: GLM-5.3's API is live now via the GLM Coding Plan. If your team uses AI-assisted vulnerability discovery, the CyberGym score justifies evaluation before the open weights arrive. You'll want an independent read on Z.ai's 2,436 claim.
- If you follow open-source AI policy: The Z.ai hold is a data point in the emerging debate about whether labs should self-certify on cyber capability before releasing weights. Watch whether August 28 holds and whether Z.ai publishes methodology documentation alongside the checkpoint.
- If you're a regular user: You won't interact with GLM-5.3 directly. What it means for you is that the software you use — banking apps, hospital portals, anything that runs on code — is increasingly likely to be scanned by AI-powered tools. That's true for defenders and, eventually, for attackers using the same technology.
Further reading
- Axios — Z.ai holds GLM-5.3 release over hacking risks — primary source on the two-week hold and Z.ai's reasoning
- VentureBeat — GLM-5.3 found a 'serious vulnerability' in Cursor — detailed coverage of what the model actually did during internal testing
- BenchLM — CyberGym leaderboard — independent scores for all evaluated models, including methodology
- D-Central — What the 84.5% score hides — skeptical read on the vulnerability count and how to interpret the numbers
- TechCrunch — OpenAI's cyber model launch context — useful for understanding where GLM-5.3 sits relative to the closed-model cyber programs
Liked this? Get the weekly digest.
Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.
Your take
How'd I do on this one?
What did I miss?
Tell Samwise (and Sam).
Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.