An AI safety test has exposed a gap in the safeguards protecting Moonshot’s Kimi models from requests involving biological weapons and assassinations.
Mindgard, which tests AI systems for vulnerabilities, told the BBC that its researchers jailbroke Kimi K2.6 and K3 Swarm in July, getting both models to provide guidance on the subjects. The process involves feeding an AI model a chain of complex instructions to test whether it will bypass safety limits set by its developers.
The firm said it emailed Moonshot about the vulnerabilities on July 27, followed up about a week later, and published a blog on Sept. 12. Moonshot only got in touch recently, after the BBC asked the company for comment, according to Mindgard.
Moonshot told the BBC it welcomed outside scrutiny "as a key pillar for building better and safer AI" and said it was discussing the findings with Mindgard. In an email to Mindgard shared with the BBC, the Chinese developer said its model had generally shown "a high refusal rate for these types of requests" in internal evaluations.
"Once the jailbreak works, it will talk about any topic, it will even freely offer up recommendations about other topics that are also nefarious, and it will be inventive and creative," Mindgard founder Peter Garraghan told the BBC World Service programme Tech Life.
Mindgard hasn't proven the weapons and assassination guidance would actually work. It argues the guardrails should have stopped the models from engaging at all. The firm also said it is confident a jailbroken Kimi K2.6 could let hackers run code on its computing resources and connect to the internet, making it a possible launchpad for cyberattacks.
The development follows recent disclosures involving AI agents from several major developers. Google said Gemini accessed three companies during security tests in May, while Australian officials said an OpenAI agent accessed public and non-public files on a government health portal in June.
Kimi is an open-weight model, meaning users can run it on their own computing infrastructure. University of Surrey professor Alan Woodward told the BBC that such models could end up in the wrong hands but could also be used for cyber-defense.
