Zhipu GLM-5.3 beats Anthropic Mythos 5 on CyberGym benchmark

Share:
Audio Loading voice…
Zhipu GLM-5.3 beats Anthropic Mythos 5 on CyberGym benchmark

Synopsis

Zhipu's GLM-5.3 has outscored Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol on the CyberGym security benchmark — the first time a Chinese AI model has publicly claimed the top spot on a standardised cyber-defence evaluation — though it trails both rivals sharply on ExploitBench.

Key Takeaways

Zhipu ( Z.ai ) launched GLM-5.3 on 14 August 2026 , its latest flagship AI model targeting cybersecurity applications.
GLM-5.3 scored 84.5% on CyberGym , beating Anthropic Mythos 5 ( 83.8% ) and OpenAI GPT-5.6 Sol ( 83.6% ), according to the company.
On ExploitBench , GLM-5.3 scored 54.4% , trailing Mythos 5 at 78% and GPT-5.6 Sol at 76.5% .
Real-world testing with Chinese security teams uncovered 2,436 vulnerabilities across 269 projects , of which 1,097 were rated medium to high severity.
The launch intensifies competition in AI-assisted cyber defence between Beijing -based labs and US frontier AI developers including Anthropic and OpenAI .

Beijing-based AI firm Zhipu, also known as Z.ai, launched its flagship GLM-5.3 model on 14 August 2026, claiming it surpassed Anthropic's frontier Mythos 5 model on a leading cybersecurity benchmark — marking a significant milestone in China's push to close the gap with Western AI systems in cyber defence applications.

CyberGym benchmark results

GLM-5.3 achieved a success rate of 84.5 per cent on CyberGym, a benchmark designed to evaluate whether AI models can identify and validate security flaws directly from source code. That figure edged out Anthropic's Mythos 5 at 83.8 per cent and OpenAI's GPT-5.6 Sol at 83.6 per cent, according to the company.

The result positions Zhipu as the first Chinese AI developer to publicly claim a top-tier ranking on a standardised cybersecurity AI evaluation, a domain that has been dominated by US-based frontier labs.

Where GLM-5.3 still trails Western rivals

The model did not match its foreign counterparts on ExploitBench, which measures how effectively AI models can progress through active exploitation chains. GLM-5.3 scored 54.4 per cent, well behind Mythos 5's 78 per cent and GPT-5.6 Sol's 76.5 per cent.

This gap suggests that while GLM-5.3 is competitive at the vulnerability-detection layer, its capability to autonomously escalate and exploit identified flaws remains meaningfully below frontier Western models — a distinction that carries real-world implications for both offensive and defensive security deployments.

Real-world testing with Chinese security teams

According to the company, Zhipu validated GLM-5.3 alongside security teams in China against live codebases, surfacing 2,436 vulnerabilities across 269 projects after expert review. Of those, 1,097 were rated medium to high severity, the company said.

The scale of the validation exercise — spanning hundreds of real-world projects — is intended to distinguish GLM-5.3's claims from purely synthetic benchmark performance, though independent third-party verification has not been disclosed.

The competitive backdrop

China's AI sector has accelerated investment in security-focused models as geopolitical tensions intensify scrutiny of critical infrastructure vulnerabilities. Firms including 360 Security Technology and units of Alibaba Group Holding have also been active in this space, while models distributed through platforms such as Hugging Face have expanded the accessibility of Chinese AI research globally.

Zhipu's claim arrives as Western governments and enterprises are actively evaluating AI-assisted security tooling, making benchmark leadership — even narrow — a commercially and diplomatically significant signal.

What's next

The durability of GLM-5.3's CyberGym lead will depend on how quickly Anthropic and OpenAI iterate, and whether independent evaluators can replicate Zhipu's numbers. The ExploitBench deficit is the clearest near-term target for the Beijing lab to close if it wants to compete for enterprise security contracts outside China.

Point of View

Not yet a full-spectrum offensive AI — a distinction that matters enormously for government and enterprise buyers. What mainstream coverage underplays is that benchmark leadership in security AI is also a soft-power signal, telling allied and non-aligned governments that Chinese models can compete at the frontier in sensitive domains. The real test will come when independent red teams attempt to replicate Zhipu's 2,436-vulnerability claim on open codebases — without that verification, the numbers remain self-reported. If the gap on ExploitBench narrows in the next model generation, the calculus for Western security vendors and policymakers changes significantly.
NationPress
14 Aug 2026

Frequently Asked Questions

What is Zhipu GLM-5.3 and what did it achieve?
GLM-5.3 is the latest flagship AI model from Beijing -based Zhipu ( Z.ai ), launched on 14 August 2026 . It achieved a 84.5% success rate on the CyberGym cybersecurity benchmark, narrowly surpassing Anthropic 's Mythos 5 ( 83.8% ) and OpenAI 's GPT-5.6 Sol ( 83.6% ), according to the company.
What is CyberGym and why does it matter for AI?
CyberGym is a benchmark that measures whether AI models can identify and validate security flaws from source code. It matters because it provides a standardised way to compare AI models' ability to assist in real-world cyber defence, making it a key metric for security teams evaluating AI-assisted vulnerability detection tools.
How does GLM-5.3 compare to GPT-5.6 Sol and Mythos 5 on ExploitBench?
GLM-5.3 scored 54.4% on ExploitBench , significantly behind Anthropic 's Mythos 5 at 78% and OpenAI 's GPT-5.6 Sol at 76.5% . ExploitBench measures how far an AI model can progress through active exploitation chains, meaning GLM-5.3 lags its Western rivals in autonomous offensive security capability.
What real-world security testing did Zhipu conduct with GLM-5.3?
According to the company, Zhipu tested GLM-5.3 with security teams in China against real-world codebases, identifying 2,436 vulnerabilities across 269 projects after expert review. Of those, 1,097 were rated medium to high severity.
Who are Zhipu's main competitors in AI cybersecurity?
Zhipu 's primary benchmark competitors are Anthropic (maker of Mythos 5 ) and OpenAI (maker of GPT-5.6 Sol ). Within China , firms such as 360 Security Technology and units of Alibaba Group Holding are also active in security-focused AI development.
Nation Press
The Trail

Connected Dots

Tracing the thread behind this story — newest first.

8 Dots
  1. Latest 3 weeks ago
  2. 3 weeks ago
  3. 3 weeks ago
  4. 1 month ago
  5. 1 month ago
  6. 1 month ago
  7. 1 month ago
  8. 2 months ago
Google Prefer NP
On Google