Tech

OpenAI, AI models out of control: platform hacked

The US giant: ‘Unprecedented cyberattack’: a popular platform for programmers has been hacked independently

FILE PHOTO: OpenAI logo is seen in this illustration taken June 11, 2026. REUTERS/Dado Ruvic/File Photo REUTERS

3' min read

Translated by AI
Versione italiana

3' min read

Translated by AI
Versione italiana

A scene straight out of a film about a digital apocalypse – something that was perhaps bound to happen sooner or later. An artificial intelligence platform that, during a drill, decides independently how to act, finds a vulnerability, escapes from a room and then spends days attacking another platform. Well, it’s happened. And the platform in question is the best-known in the Western hemisphere: ChatGPT.

OpenAI, the company behind the famous chatbot, has announced in the last few hours that its advanced artificial intelligence models spiralled out of control during security tests, autonomously breaching a popular platform for programmers. The San Francisco-based company described the incident as an ‘unprecedented cyber incident’ and announced that it will conduct a joint investigation with Hugging Face, the code and AI model-sharing platform that was the victim of the cyber attack.

Loading...

Altman’s company stated that the incident involved a combination of models, including the recent GPT-5.6 Sol and an ‘even more capable’ model currently under development. The company was seeking to assess the models’ hacking capabilities by assigning them tasks within a strictly controlled digital test environment. ‘Whilst operating in our test environment, the models devoted a significant amount of computing power to finding a way to gain unrestricted access to the internet, in order to solve the evaluation problem.’

What happened

But let’s try to understand what happened, not least because this event marks a significant milestone in AI security and cybersecurity.

Engineers at the company led by Sam Altman were conducting an internal test on the most advanced models, including one that had not yet been released to the public. It was a routine test, the sort used to gauge how well the new models perform in identifying vulnerabilities, resolving hacking issues and other similar situations.

The most important part of the story, however, is something else. Because, as is standard practice, the models had been placed in a sort of digital isolation chamber – a sandbox – to run these tests. In other words: in an isolated environment with no access to the internet. And since the test specifically concerned cyber capabilities, the security filters that normally prevent the model from carrying out hacking had been relaxed.

And this is where the scenario straight out of an American film played out: the test subject, in order to achieve a better score, decides on his own that he needs the internet. He finds a previously undiscovered flaw (a zero-day vulnerability) in the system supplying him with software packages, exploits it, and escapes from the locked room. At that point, he was free to roam the network.

Once outside, it gains access to Hugging Face’s systems – a platform that hosts open-source AI models. A target, amongst other things, that does not seem to have been chosen at random, as it is perhaps the best one for carrying out its task. And it doesn’t just gain access. Over the course of a weekend, it carries out tens of thousands of automated actions, escalates its privileges and moves through the company’s internal infrastructure.

Amidst all this, Hugging Face detects the intrusion but cannot identify who is behind it: it only knows that it is an autonomous AI agent of unprecedented sophistication. And it makes this public.

A dramatic turn of events in the last few hours: OpenAI has admitted that the mysterious attacker was one of its own models, which had gone rogue during testing. No human hacker gave the command to attack Hugging Face. The model did it all on its own, as a side effect of the objective it had been given (“solve this test well”).

This incident is almost certainly the first documented case in which an AI has evaded its creator’s containment measures and, acting on its own, compromised the infrastructure of a real third-party company. And perhaps it demonstrates that state-of-the-art models now possess offensive capabilities comparable to (or perhaps even superior to) those of elite hackers, and that they can use them in unexpected ways to achieve a goal. This is precisely the scenario that AI security researchers have been discussing for years as a theoretical risk. A risk that is now no longer theoretical.

Copyright reserved ©
Loading...

Brand connect

Loading...

Newsletter

Notizie e approfondimenti sugli avvenimenti politici, economici e finanziari.

Iscriviti