Anthropic admits that its AI breached the systems of three companies during a test
It happened during the testing phase. The incident came just a few days after OpenAI’s warning
Anthropic has stated that its artificial intelligence models breached the systems of three other organisations during testing, just days after OpenAI, the company behind ChatGPT, had raised concerns about AI control following the revelation that its unauthorised models had breached another company’s systems.
The case
Anthropic, the San Francisco-based AI company that created Claude, announced on its website on Thursday that it had discovered the three incidents after examining over 141,000 evaluation runs.
In response to the OpenAI incident, Anthropic stated that it had launched a “large-scale” cybersecurity review specifically aimed at verifying whether its AI models had been able to access the internet within test environments that were supposed to be isolated.
Anthropic has clarified that the models involved in the incidents are Claude Opus 4.7, Claude Mythos 5 and an internal test model used for research. The first incidents date back to April, the AI company said.
“Claude compromised the infrastructure of the affected organisations using basic techniques,” said Anthropic, such as exploiting weak passwords.
