Skip to main content
Back to Newswire
Security

UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities

Performance of Kimi K3 and other models on an exploit development benchmark (ExploitBench). Higher success rate indicates greater cyber capability. Error bars represent 95% confidence intervals. ExploitBench measures the capability of a model to develop end-to-end exploits given a vulnerability. Image: Primary
The UK Artificial Intelligence Security Institute and the U.S. Center for AI Standards and Innovation said Wednesday they conducted a joint preliminary evaluation of Moonshot AI's Kimi K3 model, which was released July 16 and slated for open-weight release by July 27. The assessment focused on cyber capabilities using the ExploitBench benchmark and a simulated corporate network attack range called "The Last Ones." Kimi K3 achieved a 32% success rate on ExploitBench, outperforming GLM-5.2, the most cyber-capable open-weight model as of June 2026, which scored 24%. However, Kimi K3 failed to achieve arbitrary code execution on any of the 41 ExploitBench tasks, while the most capable models achieved it on 20 of 41 samples on average. On the 32-step simulated attack range, Kimi K3 reached step 17 on average compared to 28.5 steps for leading U.S. models. Kimi K3 completed the full range in one of 10 attempts within the token limit, indicating it can autonomously attack small, weakly defended enterprise systems when directed and given initial network access. The institutes noted the results are preliminary and based on a selective set of evaluations due to the model's hosting setup.
Sources
In this story
Published by Tech & Business, a media brand covering technology and business. This story was sourced from nist.gov and reviewed by the T&B editorial agent team.