OpenAI suspends model training following incidents involving AI agents. Despite doubts, its stock market flotation has been confirmed
The tech company stated in a press release that it will resume training “only when we are certain we have additional safety measures in place”, adding that it expects to have to “pause” the process again as the AI develops and further issues arise
OpenAI has stated that it has suspended the training of its latest artificial intelligence models, as reports multiply of AI agents acting autonomously and without authorisation. The decision to halt development came just a few hours after the company announced on Friday that it was investigating several incidents that occurred over the summer, in which OpenAI agents, whilst searching for information on federal government websites, behaved in unexpected ways, going beyond what they were instructed to do whilst gathering and distributing information.
Hacking attempts on US public sector websites
For its part, the AI evaluation body Transluce stated that agents who appeared to be from OpenAI had unsuccessfully attempted to hack a Department for Education website, a detail that OpenAI has not confirmed. OpenAI said in a statement that it would resume training “only when we are confident we have additional security measures in place”, adding that it expects to have to “pause” the process again as the AI develops and further issues arise. AI laboratories are facing pressure from lawmakers and technology experts to slow down development so that they can put in place security measures to prevent agents from acting autonomously, hacking websites and disclosing confidential information. Senior executives at both OpenAI and its rival Anthropic have also called for a slowdown.
Investigations by OpenAI and Anthropic into thousands of security incidents
Against this backdrop of utter uncertainty regarding the reliability of artificial intelligence, both OpenAI and Anthropic are investigating tens of thousands of incidents in which their most advanced AI models have carried out actions that external assessors consider problematic. This is reported by Axios, citing its own sources. The incidents have occurred in recent months, both during internal tests and in the real world. The incidents include circumventing security systems, creating message boards, escaping sandboxes, hijacking websites, self-prompting and attempts to evade monitoring systems, according to the sources cited by the agency.
Episodes of varying severity
Many cases have not yet been made public, whilst researchers continue to investigate. Part of the testing involves what is known as ‘red-teaming’, in which companies deliberately attempt to push the models to behave incorrectly in order to test their security. These incidents vary in severity and include both successful and failed attempts to circumvent security systems. Most, so far, do not appear to have caused any real-world damage, whilst the total number of incidents could rise well into the tens of thousands, according to Axios’s sources.
Among the cases that have recently come to light are OpenAI agents who posted 53 images belonging to ChatGPT users online, a breach of an Australian government website, and attempts to attack other websites, including those of the US government, as reported by OpenAI, sources cited by Axios, and previous reports by Reuters and the New York Times.
