Kimi K3 Cyber Capabilities Assessed: Significant Gap with US Models

Kimi K3 Cyber Capabilities Assessed: Significant Gap with US Models

First seen 24 Jul 2026, 15:52 UTC Nistwww.aisi.gov.ukScmp 86% similarity 43.0

Article Content

Browse articles
ThreatCluster

A joint evaluation by the UK AISI and US CAISI assessed China's Kimi K3 AI model, revealing it has a cyber capability score of 32.2%, significantly lower than top US models averaging 76.2%. Kimi K3, released on July 16, 2026, failed to achieve arbitrary code execution across all 41 tasks in the ExploitBench benchmark, while leading US models succeeded in 20 tasks. The evaluation raises concerns about the effectiveness of Kimi K3 in launching cyberattacks, despite outperforming a domestic competitor, Zhipu AI's GLM-5.2, which scored 24.4%. The findings challenge perceptions of the rapid advancement of Chinese AI in cybersecurity. The report highlights the need for further assessments as Kimi K3 is set for open-weight release on July 27, 2026.

Key Points: • Kimi K3 scored 32.2% on the ExploitBench benchmark, far below US rivals. • The model failed to achieve arbitrary code execution in all tested tasks. • The assessment indicates a significant gap in cyber capabilities between US and Chinese AI models.

ThreatCluster AI

Timeline

2026-07-16
Kimi K3 released
Moonshot AI launched Kimi K3, its latest AI model, focusing on cyber capabilities.
Nist
2026-07-23
Kimi K3 evaluation published
UK AISI and US CAISI released findings showing Kimi K3's cyber capabilities lag behind US models.
Nist
2026-07-24
SCMP reports on Kimi K3's performance
SCMP highlighted Kimi K3's 32.2% score and its inability to execute high-level exploits.
Scmp

Community

Browse all →