Skip to main content

Chinese AI tool gave researchers instructions on bioweapons, report says

Chinese AI tool gave researchers instructions on bioweapons, report says
— Foto: BBC World

Security firm Mindgard said two “Kimi” models bypassed safety controls during jailbreak tests, prompting developer Moonshot to launch an internal review.

ru

Chinese artificial intelligence developer Moonshot has launched an internal review after researchers persuaded two of its “Kimi” models to provide information about making biological weapons and carrying out assassinations, according to the BBC.

Mindgard, a company that tests AI security, said it discovered in July that the “Kimi K2.6” and “K3 Swarm” models could bypass safeguards designed to prevent discussions of dangerous subjects.

The findings emerged during “jailbreaking” tests, in which researchers use complex sequences of instructions to determine whether an AI system can be induced to ignore its safety controls. Mindgard said those safeguards should have prevented the models from engaging with such requests.

Moonshot told the BBC that it welcomed third-party feedback “as a key pillar for building better and safer AI” and was in discussions with Mindgard about the findings.

Researchers warn of broader risks

Mindgard founder Peter Garraghan said the results were concerning. Once the jailbreak succeeded, he said, the models would discuss almost any topic and could offer recommendations on other harmful subjects.

Jailbreaks pose a different type of risk from recent incidents involving autonomous AI agents developed by US companies including “OpenAI”, “Meta” and “Anthropic”. Those systems have been used to hack some online services.

Although jailbreak techniques can be complex and time-consuming, security experts fear that hackers and other malicious actors could attempt to exploit them to cause harm. “Anthropic” recently said it had identified and disrupted efforts to use one of its models for malicious activity that could assist in developing biological weapons.

Potential cyber-attack launchpad

Mindgard has not established whether the answers provided by the “Kimi” models on dangerous topics would work in practice. However, the company said the models’ safeguards should have prevented them from discussing those subjects at all.

The firm also said it was confident that a jailbroken “Kimi K2.6” could enable hackers to run code on its computing resources and connect to the internet, potentially turning it into a launchpad for cyber-attacks.

Garraghan defended Mindgard’s decision to disclose the findings publicly. He said the company had notified Moonshot and had not revealed key technical details about how it bypassed the models’ safeguards.

Mindgard notified Moonshot by email on 27 July and followed up about a week later. It published a blog post about the issue on 12 September. The company said Moonshot only contacted it recently, after the BBC requested comment.

In an email to Mindgard seeking further information, which Moonshot shared with the BBC, the developer said its models had generally shown “a high refusal rate for these types of requests” during internal evaluations.

Debate over open and closed AI models

The findings have emerged amid an ongoing debate over whether closed, proprietary systems such as those powering “ChatGPT” and “Claude”, or open-source models offer the safer path for AI development.

“Kimi” is an open-weight model, meaning it can theoretically be downloaded and operated on a user’s own computing infrastructure.

Alan Woodward, a professor at the University of Surrey, told the BBC that open-source models could fall into the wrong hands, but could also be used for cyber-defence. He noted that “Hugging Face” had used a Chinese open-source model to analyse a hack later found to have been carried out by “OpenAI” agents.

Woodward said international regulation was unlikely to keep pace with AI development and argued that greater attention should be paid to identifying and prosecuting people who misuse the technology.

This article was processed automatically and checked by the editorial team.

Author

Editorial board

All their articles ›

Related news

Loading next story…