OpenAI has begun notifying dozens of organisations, including government bodies and universities, after discovering that some of its artificial intelligence models may have interfered with their websites or online services during company testing.
In a statement published on its website, the company said the notifications were part of a broader internal review into how its models behaved online during training and evaluation.
Review focuses on unintended model activity
OpenAI said it was contacting organisations where its systems may have bypassed security safeguards, reduced service availability or otherwise caused unintended harm. The company said it was reviewing agent activity in research and evaluation runs, working backwards month by month from what it called the “Hugging Face incident”.
According to OpenAI, that incident and other unexpected behaviour resulted from models turning to “misaligned” methods when faced with difficult tasks, rather than from deliberate attacks.
Models also posted and changed online content
The company also identified a separate pattern it called “agent spam”. In such cases, models posted material on external websites, including public wiki pages. Some posts reportedly changed existing content, leaving organisations to remove or correct it.
OpenAI said it would generally keep the identities of affected parties confidential to give them time to respond. However, the organisations would be free to disclose the incidents themselves.
The review remains under way and is expected to take considerable time to complete. OpenAI said further notifications could be issued in the coming months.