Kimi K3 scores 32.2% on cyber benchmark, far below US AI rivals
Synopsis
Key Takeaways
Moonshot AI's Kimi K3, widely regarded as China's most capable large language model, scored 32.2 per cent on a standardised cybersecurity exploit benchmark — less than half the 76.2 per cent average posted by leading US models — according to a joint government study published on Thursday, 24 July 2026. The findings directly challenge concerns in Washington about the pace at which Chinese open-source AI is closing the gap with American frontier models.
Who conducted the research
The evaluation was carried out jointly by the UK Artificial Intelligence Security Institute (AISI), a research arm of the UK Department for Science, Innovation and Technology, and the US Centre for AI Standards and Innovation (CAISI), which operates under the US Department of Commerce's National Institute of Standards and Technology. The two bodies used ExploitBench, a publicly available benchmark designed to measure an AI model's ability to identify and develop exploits targeting cybersecurity vulnerabilities.
What the benchmark revealed
Kimi K3 achieved an overall ExploitBench score of 32.2 per cent, placing it ahead of domestic rival Zhipu AI's GLM-5.2, which scored 24.4 per cent. However, the gap with unnamed leading US models — which averaged 76.2 per cent — was substantial. The report described Kimi K3's performance as 'significantly below the most recent frontier cyber-capable models.'
Why it matters: the arbitrary code execution gap
A particularly telling result was Kimi K3's failure to achieve arbitrary code execution across all 41 ExploitBench tasks. Arbitrary code execution represents the highest-level exploit category, granting an attacker full control of a target system. By contrast, the leading US models achieved this on 20 of the 41 tasks — underscoring a meaningful capability divide at the most dangerous end of the spectrum.
The competitive backdrop
The study arrives as policymakers in both Washington and London have escalated scrutiny of Chinese AI models, particularly open-source releases from companies such as Moonshot AI and Zhipu AI. Proponents of tighter export controls have argued that capable Chinese models could be adapted for offensive cyber operations. The AISI/CAISI findings suggest that, at least on measurable exploit-generation tasks, that risk remains materially lower than for frontier US systems — which themselves post scores that regulators may find troubling.
What's next
The report does not assess newer or unreleased Chinese models, and the benchmark landscape is evolving rapidly. As Moonshot AI and its peers continue iterating, the gap flagged in this study will be a key metric for both security researchers and policymakers to track. The question of whether open-source model releases accelerate the closing of this divide will likely shape the next round of US AI export-control deliberations.