Anthropic’s AI also affects external services during testing
Following the OpenAI-Hugging Face case, Anthropic has now admitted its involvement. But it can’t always be the AI’s fault.
Anthropic has revealed that its Claude AI models hacked three organisations whilst the start-up was testing their cyber capabilities, a week after OpenAI had reported a similar incident.
In practice, Claude is said to have gained unauthorised access to external organisations whilst carrying out an assessment of his cyber-offensive tasks. “A misunderstanding” gave Claude access to the internet in his test environment, when it should have been blocked, Anthropic said.
The admission comes a week after rival OpenAI admitted that two of its models had hacked Hugging Face whilst the model’s developer was testing its technology this month. The models escaped from their test environment due to a software vulnerability, allowing them to access the internet and carry out the cyberattack.
As with OpenAI, the news has reignited a debate that is set to resurface with increasing frequency in the coming years: can we say that AI itself is to blame? The short answer is no, and it is worth understanding why.
It is wrong to talk of attacks devised by AI models without any input from humans, and it is time to ask what mistakes were made – whether intentionally or not – by those who tested the models responsible for the attacks.

