Kimi K3 scores 32.2% on cyber benchmark, far below US AI rivals

Share:
Audio Loading voice…
Kimi K3 scores 32.2% on cyber benchmark, far below US AI rivals

Synopsis

A joint UK-US government study found China's most powerful AI model, Kimi K3, scored just 32.2% on a cyberattack capability benchmark — versus 76.2% for leading US models — and failed every arbitrary code execution task, directly undercutting fears of imminent Chinese AI cyber parity.

Key Takeaways

Kimi K3 , developed by Moonshot AI , scored 32.2% on the ExploitBench cybersecurity benchmark, per a report published on 24 July 2026 .
Leading unnamed US AI models averaged 76.2% on the same benchmark, a gap of more than 44 percentage points .
Kimi K3 failed to achieve arbitrary code execution on any of the 41 ExploitBench tasks; top US models succeeded on 20 .
Domestic rival Zhipu AI 's GLM-5.2 scored 24.4% , placing it below Kimi K3 but far behind US frontier models.
The study was conducted jointly by the UK AISI and the US CAISI , both government AI safety bodies.

Moonshot AI's Kimi K3, widely regarded as China's most capable large language model, scored 32.2 per cent on a standardised cybersecurity exploit benchmark — less than half the 76.2 per cent average posted by leading US models — according to a joint government study published on Thursday, 24 July 2026. The findings directly challenge concerns in Washington about the pace at which Chinese open-source AI is closing the gap with American frontier models.

Who conducted the research

The evaluation was carried out jointly by the UK Artificial Intelligence Security Institute (AISI), a research arm of the UK Department for Science, Innovation and Technology, and the US Centre for AI Standards and Innovation (CAISI), which operates under the US Department of Commerce's National Institute of Standards and Technology. The two bodies used ExploitBench, a publicly available benchmark designed to measure an AI model's ability to identify and develop exploits targeting cybersecurity vulnerabilities.

What the benchmark revealed

Kimi K3 achieved an overall ExploitBench score of 32.2 per cent, placing it ahead of domestic rival Zhipu AI's GLM-5.2, which scored 24.4 per cent. However, the gap with unnamed leading US models — which averaged 76.2 per cent — was substantial. The report described Kimi K3's performance as 'significantly below the most recent frontier cyber-capable models.'

Why it matters: the arbitrary code execution gap

A particularly telling result was Kimi K3's failure to achieve arbitrary code execution across all 41 ExploitBench tasks. Arbitrary code execution represents the highest-level exploit category, granting an attacker full control of a target system. By contrast, the leading US models achieved this on 20 of the 41 tasks — underscoring a meaningful capability divide at the most dangerous end of the spectrum.

The competitive backdrop

The study arrives as policymakers in both Washington and London have escalated scrutiny of Chinese AI models, particularly open-source releases from companies such as Moonshot AI and Zhipu AI. Proponents of tighter export controls have argued that capable Chinese models could be adapted for offensive cyber operations. The AISI/CAISI findings suggest that, at least on measurable exploit-generation tasks, that risk remains materially lower than for frontier US systems — which themselves post scores that regulators may find troubling.

What's next

The report does not assess newer or unreleased Chinese models, and the benchmark landscape is evolving rapidly. As Moonshot AI and its peers continue iterating, the gap flagged in this study will be a key metric for both security researchers and policymakers to track. The question of whether open-source model releases accelerate the closing of this divide will likely shape the next round of US AI export-control deliberations.

Point of View

Not just policymakers focused on China. The study is also a snapshot: ExploitBench scores are moving targets, and the open-source release cadence from Chinese labs means the 44-percentage-point gap could narrow faster than export-control cycles can respond. Ultimately, the report may inadvertently make the stronger case for hardening critical infrastructure against all capable AI systems, regardless of origin.
NationPress
24 Jul 2026

Frequently Asked Questions

How did Kimi K3 perform on the ExploitBench cybersecurity test?
Kimi K3 scored 32.2% on ExploitBench , a public benchmark measuring an AI model's ability to generate exploits for cybersecurity vulnerabilities. This placed it ahead of Zhipu AI 's GLM-5.2 ( 24.4% ) but well below leading US models, which averaged 76.2% .
Who published the Kimi K3 cybersecurity study?
The report was published jointly on 24 July 2026 by the UK Artificial Intelligence Security Institute (AISI) and the US Centre for AI Standards and Innovation (CAISI) . The AISI sits within the UK Department for Science, Innovation and Technology , while CAISI operates under the US Department of Commerce 's National Institute of Standards and Technology .
What is arbitrary code execution and why does it matter for AI safety?
Arbitrary code execution is the highest-level exploit category, allowing an attacker to gain full control of a target system. Kimi K3 failed to achieve it on all 41 ExploitBench tasks, while top US models succeeded on 20 — highlighting a significant capability gap at the most dangerous end of the cyberattack spectrum.
Does the study mean Chinese AI poses no cyber threat?
The study measures a specific, benchmarked capability at a single point in time and does not make a broader threat assessment. Kimi K3 's 32.2% score is 'significantly below' frontier models, according to the report, but the open-source nature of Chinese AI development means capabilities could improve rapidly in future model iterations.
How does Kimi K3 compare to other Chinese AI models?
Kimi K3 outperformed domestic rival Zhipu AI 's GLM-5.2 , which scored 24.4% on ExploitBench . Both Chinese models, however, trail the unnamed leading US models by a wide margin, with the top US systems averaging more than double Kimi K3 's score.
Nation Press
The Trail

Connected Dots

Tracing the thread behind this story — newest first.

8 Dots
  1. Latest 13 hours ago
  2. 20 hours ago
  3. 2 days ago
  4. 2 days ago
  5. 3 days ago
  6. 4 days ago
  7. 6 days ago
  8. 1 week ago
Google Prefer NP
On Google