OpenAI suspends model training following incidents involving AI agents. IPO confirmed despite doubts
The tech company stated in a press release that it would resume training “only when we are certain that we have additional safety measures in place”, adding that it expects to have to “pause” the process again as the AI develops and further issues arise
OpenAI has stated that it has suspended the training of its latest artificial intelligence models, as reports multiply of AI agents acting autonomously and without authorisation. The decision to halt development came just a few hours after the company announced on Friday that it was investigating several incidents that occurred over the summer, in which OpenAI agents, whilst searching for information on federal government websites, behaved in unexpected ways, going beyond what was required of them whilst gathering and distributing information.
Attempted hacking of US public sector websites
For its part, the AI evaluation body Transluce stated that agents who appeared to be from OpenAI had unsuccessfully attempted to hack a Department for Education website – a detail that OpenAI has not confirmed. OpenAI said in a statement that it would resume training “only when we are confident we have additional security measures in place”, adding that it expects to have to “pause” the process again as the AI develops and further issues arise. AI laboratories are facing pressure from legislators and technology sector experts to slow down development so that they can put in place security measures to prevent agents from acting autonomously, hacking websites and disclosing confidential information. Senior executives at both OpenAI and its rival Anthropic have also called for a slowdown.
Investigations by OpenAI and Anthropic into thousands of security incidents
Against this backdrop of utter uncertainty regarding the reliability of artificial intelligence, both OpenAI and Anthropic are investigating tens of thousands of incidents in which their most advanced AI models have carried out actions that external assessors consider problematic. Axios reports this, citing its own sources. The incidents have occurred in recent months, both during internal testing and in the real world. The incidents include circumventing security systems, creating message boards, escaping from sandboxes, hijacking websites, self-prompting and attempts to evade monitoring systems, according to sources cited by the agency.
Episodes of varying severity
Many cases have not yet been made public, whilst researchers continue to investigate. Part of the testing involves so-called ‘red-teaming’, in which companies deliberately attempt to push the models to behave incorrectly in order to test their security. These incidents vary in severity and include both successful and failed attempts to circumvent security systems. Most, so far, do not appear to have caused any real-world damage, whilst the total number of incidents could rise well into the tens of thousands, according to Axios’ sources.
Among the cases that have recently come to light are OpenAI agents who posted 53 images belonging to ChatGPT users online, a breach of an Australian government website and attempts to attack other websites, including those of the US government, as reported by OpenAI, sources cited by Axios, and previous reports by Reuters and the New York Times.
