Zhipu GLM-5.3 beats Anthropic Mythos 5 on CyberGym benchmark
Synopsis
Key Takeaways
Beijing-based AI firm Zhipu, also known as Z.ai, launched its flagship GLM-5.3 model on 14 August 2026, claiming it surpassed Anthropic's frontier Mythos 5 model on a leading cybersecurity benchmark — marking a significant milestone in China's push to close the gap with Western AI systems in cyber defence applications.
CyberGym benchmark results
GLM-5.3 achieved a success rate of 84.5 per cent on CyberGym, a benchmark designed to evaluate whether AI models can identify and validate security flaws directly from source code. That figure edged out Anthropic's Mythos 5 at 83.8 per cent and OpenAI's GPT-5.6 Sol at 83.6 per cent, according to the company.
The result positions Zhipu as the first Chinese AI developer to publicly claim a top-tier ranking on a standardised cybersecurity AI evaluation, a domain that has been dominated by US-based frontier labs.
Where GLM-5.3 still trails Western rivals
The model did not match its foreign counterparts on ExploitBench, which measures how effectively AI models can progress through active exploitation chains. GLM-5.3 scored 54.4 per cent, well behind Mythos 5's 78 per cent and GPT-5.6 Sol's 76.5 per cent.
This gap suggests that while GLM-5.3 is competitive at the vulnerability-detection layer, its capability to autonomously escalate and exploit identified flaws remains meaningfully below frontier Western models — a distinction that carries real-world implications for both offensive and defensive security deployments.
Real-world testing with Chinese security teams
According to the company, Zhipu validated GLM-5.3 alongside security teams in China against live codebases, surfacing 2,436 vulnerabilities across 269 projects after expert review. Of those, 1,097 were rated medium to high severity, the company said.
The scale of the validation exercise — spanning hundreds of real-world projects — is intended to distinguish GLM-5.3's claims from purely synthetic benchmark performance, though independent third-party verification has not been disclosed.
The competitive backdrop
China's AI sector has accelerated investment in security-focused models as geopolitical tensions intensify scrutiny of critical infrastructure vulnerabilities. Firms including 360 Security Technology and units of Alibaba Group Holding have also been active in this space, while models distributed through platforms such as Hugging Face have expanded the accessibility of Chinese AI research globally.
Zhipu's claim arrives as Western governments and enterprises are actively evaluating AI-assisted security tooling, making benchmark leadership — even narrow — a commercially and diplomatically significant signal.
What's next
The durability of GLM-5.3's CyberGym lead will depend on how quickly Anthropic and OpenAI iterate, and whether independent evaluators can replicate Zhipu's numbers. The ExploitBench deficit is the clearest near-term target for the Beijing lab to close if it wants to compete for enterprise security contracts outside China.