Kimi K3 scores 32% on cyber exploit benchmark as US models hit 76%, raising distillation questions
A joint evaluation by UK AISI and US CAISI found Moonshot AI's Kimi K3 scored 32.2% on ExploitBench versus 76.2% for leading US models, while its safeguards failed to block exploit development or simulated network attacks.
The British AI Security Institute and the U.S. Center for AI Standards and Innovation jointly evaluated Moonshot AI's Kimi K3 on offensive cyber tasks, finding it scored 32.2 percent on the ExploitBench benchmark compared with 76.2 percent for leading U.S. models. The model's safeguards did not block exploit development or simulated attacks.
Kimi K3 reached step 17 out of 32 on average in the "The Last Ones" simulated corporate network attack, while leading U.S. models reached 28.5 steps. The model completed the full 32-step attack path in one of ten attempts within a 100 million token limit, showing unreliable but present capability.
ExploitBench results show wide gap on critical vulnerabilities
The institutes used Carnegie Mellon University's ExploitBench, which tests exploit development against 41 vulnerabilities found in Chrome's V8 engine after 2023. Kimi K3 did not achieve Arbitrary Code Execution on any of the 41 tasks, while leading U.S. models achieved ACE in 20 tasks. China's GLM-5.2 scored 24.4 percent.
- Kimi K3 scored 32.2 percent on ExploitBench versus 76.2 percent for leading U.S. models
- The model failed to reach Arbitrary Code Execution on any of 41 Chrome V8 vulnerabilities
- In the TLO network attack simulation, Kimi K3 averaged step 17 of 32 versus 28.5 for U.S. models
- GLM-5.2 scored 24.4 percent on ExploitBench and averaged step 11 in TLO
- Kimi K3 completed the full TLO attack path in one of ten attempts
Distillation allegations gain support from benchmark gap
The gap between Kimi K3's strong general benchmark scores and weaker cyber performance aligns with allegations that Moonshot AI distilled Anthropic's models. U.S. science advisor Michael Kratsios accused Moonshot AI of using Anthropic's Fable outputs as training data. One explanation is that Kimi K3 was trained mostly on Claude outputs covering general knowledge and programming, but Anthropic's safety classifiers specifically block advanced cyber capabilities.
The institutes tested U.S. closed-weight models with system-level safeguards disabled to measure maximum capabilities. Those safeguards remain enabled in publicly available versions. A time-series analysis by CAISI shows Chinese models gaining cyber capabilities since early 2025 but consistently trailing U.S. counterparts by four to seven months for open models.
AISI warns that the growing cyber capabilities of open models create "a persistent and irreversible risk of misuse." The evaluation did not account for active defense in the TLO simulation, meaning real-world scenarios could differ. Redis shipped seven security releases on July 23 after researchers published authenticated RCE proof-of-concepts for stock Redis versions, though those findings were separate from the Kimi K3 evaluation.
Fact check
-
Kimi K3 scored 32.2 percent on the ExploitBench benchmark
verified · source
-
Leading US models scored 76.2 percent on ExploitBench
verified · source
-
Kimi K3 did not achieve Arbitrary Code Execution on any of 41 Chrome V8 vulnerabilities
verified · source
-
Kimi K3 reached step 17 out of 32 on average in the TLO network attack simulation
verified · source
-
Redis shipped seven security releases on July 23 after researchers published authenticated RCE PoCs
reported · source
Source reporting (2)
Join the conversation
You need to be registered and logged in to comment on blog articles.
0 Comments
No comments yet
Be the first to share your thoughts on this article.