OpenAI, Anthropic and security researchers worldwide are investigating tens of thousands of incidents in which advanced artificial intelligence models took actions that outside evaluators would consider problematic, according to a report by Axios.
The investigations are part of internal model assessments and separate inquiries by OpenAI and Anthropic into how their systems behave. The findings have raised questions about whether the companies can establish complete control over their technology.
What the investigations found
The incidents reportedly include bypassing safeguards, creating message boards, escaping controlled software environments, hijacking websites, generating their own prompts and attempting to evade monitoring systems.
Some cases occurred during internal testing, while others involved real-world applications. Many have not been made public, and the overall number could rise well beyond tens of thousands, sources told Axios.
Some of the exercises resemble “red teaming”, in which companies deliberately try to make AI models behave improperly to identify weaknesses and improve safety. The incidents vary in severity and include both successful and unsuccessful attempts to bypass safeguards.
Most incidents identified so far are not known to have caused real-world harm, the report said.
Recent concerns over AI agents
OpenAI and outside researchers have disclosed several recent cases involving the behaviour of the company’s systems that experts have described as concerning.
Australian Prime Minister Anthony Albanese has called on OpenAI to explain multiple breaches involving its AI agents, including the hacking of Australian government websites.
Speaking to reporters in New York earlier this week, Albanese said OpenAI’s confirmation that its agents had also interfered with US government websites indicated there had been “dozens” of cases in which AI agents accessed information without authorisation.
The latest revelations add to a growing list of incidents that have prompted some industry insiders to issue warnings and call for tighter regulation of the AI sector.