Next upHack for Humanity: San Francisco (powered by Google Gemini)
News

Z.ai's open-weight GLM-5.2 matches Claude Opus 4.8 on cybersecurity benchmarks at lower cost

Two independent evaluators found Z.ai's open-weight GLM-5.2 matching Anthropic's Opus 4.8 on cybersecurity tasks at roughly 2.2 times lower API cost.

Dmytro Spodarets
Jun 28, 2026 · 1 min read

Two independent evaluations published June 23 to 25, 2026 found that Z.ai’s open-weight GLM-5.2 matches Anthropic’s Claude Opus 4.8 on cybersecurity tasks while costing roughly 2.2 times less to run. The results land days after a US export ban cut off Anthropic’s Fable 5 and Mythos 5, sharpening a debate over whether open models blunt such controls.

GLM-5.2, a 753-billion-parameter Mixture-of-Experts model released under an MIT license by Z.ai (formerly Zhipu AI), published its weights on Hugging Face on June 16. It handles a one-million-token context window, per Z.ai’s documentation.

The benchmark numbers come from outside testers, not the vendor. Graphistry, a security analytics firm, reported on June 23 that the model tied Opus 4.7 and 4.8 on its CyBT-CTF capture-the-flag evaluation at a 28-of-59 solve rate, beating the next-best open model by 20 points. Two days later, Semgrep reported GLM-5.2 scored 39 percent F1 on detecting insecure direct object reference flaws, edging Claude Code at 32 percent, at about $0.17 per vulnerability found.

Both evaluators urged caution. Semgrep said its result reflects one task, one dataset and a single run. Graphistry floated a speculative, unconfirmed possibility of “illegal distillation” from US models, citing an 0.80 output correlation between GLM-5.2 and OpenAI and Anthropic systems against 0.63 between those two, a statistical hint, not proof, and one Z.ai has not addressed.

Vendor-reported figures put GLM-5.2 at 62.1 on SWE-bench Pro and 91.2 on GPQA-Diamond; those are the company’s own numbers. The open weights cut both ways: anyone can strip safety controls from a downloaded model, lowering the barrier to misuse even as the same openness undercuts the rationale for export curbs.

The coming weeks will show whether security teams adopt GLM-5.2 in production or treat the benchmark wins as lab artifacts.


Dmytro Spodarets
Dmytro Spodarets
Founder & Editor-in-Chief

Founder and Chief Editor of Data Phoenix — a San Francisco Bay Area media and education platform focused on AI and Data.

More news