Cybersecurity

Anthropic’s AI also affects external services during testing

Following the OpenAI-Hugging Face case, Anthropic has now admitted its involvement. But it can’t always be the AI’s fault.

 REUTERS

3' min read

Translated by AI
Versione italiana

3' min read

Translated by AI
Versione italiana

Anthropic has revealed that its Claude AI models hacked three organisations whilst the start-up was testing their cyber capabilities, a week after OpenAI had reported a similar incident.

In practice, Claude is said to have gained unauthorised access to external organisations whilst carrying out an assessment of his cyber-offensive tasks. “A misunderstanding” gave Claude access to the internet in his test environment, when it should have been blocked, Anthropic said.

Loading...

The admission comes a week after rival OpenAI admitted that two of its models had hacked Hugging Face whilst the model’s developer was testing its technology this month. The models escaped from their test environment due to a software vulnerability, allowing them to access the internet and carry out the cyberattack.

As with OpenAI, the news has reignited a debate that is set to resurface with increasing frequency in the coming years: can we say that AI itself is to blame? The short answer is no, and it is worth understanding why.

It is wrong to talk of attacks devised by AI models without any input from humans, and it is time to ask what mistakes were made – whether intentionally or not – by those who tested the models responsible for the attacks.

Headlines such as ‘Artificial intelligence has hacked a system’ are effective, but they can be misleading. A model has no intentions of its own; it does not decide independently what to do with its day. It is a tool configured by people who assign it objectives, privileges and operational limits. If the tool goes beyond those limits, the problem does not lie with the tool: it lies with those who chose those limits – or failed to choose them at all.

The latest generation of AI agents do more than just write text. They use browsers, execute code, explore digital environments and chain together actions to achieve a result. This makes them useful, and it is precisely for this reason that they are also unpredictable, as no one can foresee all the ways in which they might interpret a task. An agent that finds an unexpected shortcut has not ‘got out of control’: it is proof that the control was not tight enough to anticipate that shortcut.

Then there is a narrative issue that is worth calling by its proper name: marketing. Portraying AI as an entity capable of acting on its own, independently, reinforces the idea of an unstoppable technology. It’s a story that sells well, on social media and in newspaper headlines. In reality, these systems operate within architectures designed, monitored and managed by human beings, with all the flaws that this entails.

In cases like this, specific mistakes tend to recur. The first is granting too many privileges to a system still in the experimental phase: unrestricted internet access, real credentials, production environments, without active supervision. The second is confusing technical capability with reliability: a model that easily identifies vulnerabilities is not automatically ready to operate without supervision in a real-world context. The third is weak containment: poorly designed sandboxes, no human approval for critical actions, incomplete logs, and no quick way to halt execution if something goes wrong. The fourth, less technical but equally significant, is to report these incidents by emphasising the machine’s autonomy rather than the shortcomings of those who put it into operation.

None of these issues are actually new. The cybersecurity community has been aware of them for years under different names: misconfigured cloud environments, exposed APIs, unpatched IoT devices. Whenever a technology increases the available operational capabilities, it also increases the attack surface that someone can exploit, whether it be an external attacker or unforeseen internal behaviour. AI agents simply add another layer to an already long list.

Loading...

Incidents such as this are, predictably, set to increase. Not because AI is becoming uncontrollable, but because it will be increasingly integrated into business processes, with growing access to real-world data and tools. The correct approach is similar to that already adopted for critical infrastructure: least privilege, continuous monitoring, segregated environments, independent audits, and clearly identifiable human accountability for every automated action.

These cases do not prove that artificial intelligence has become dangerous in its own right. They demonstrate something far more unsettling for those who build it: when real operational capabilities are entrusted to increasingly powerful systems, the only variable that really matters remains the human ability to control them. The rest, for now, is just a catchy headline.

Copyright reserved ©
Loading...

Brand connect

Loading...

Newsletter

Notizie e approfondimenti sugli avvenimenti politici, economici e finanziari.

Iscriviti