Artificial intelligence

Dario Amodei of Anthropic: slowing down the development of AI to avoid catastrophic risks

New warning: ‘Potentially catastrophic damage’. Musk: ‘Dario Amodei is right’. OpenAI confirms another incident

Il CEO di Anthropic, Dario Amodei, partecipa a una sessione dell’AI Impact Summit 2026 al Bharat Mandapam di Nuova Delhi, in India, il 19 febbraio 2026. L'India ospita il vertice dal 16 al 20 febbraio 2026.  EPA/RAJAT GUPTA EPA

3' min read

Translated by AI
Versione italiana

3' min read

Translated by AI
Versione italiana

As new incidents and attacks generated autonomously by artificial intelligence come to light, the leadership at Anthropic is taking a stance on the growing security concerns. “We need to slow down the pace at which we improve the capabilities of AI models. Progress will still seem rapid, and we must make wise use of the time we gain,” writes Anthropic’s CEO, Dario Amodei, in a lengthy post in which he calls for a measured approach to the potential risks associated with AI. Yesterday, OpenAI is also reported to have urged its staff to slow down the development of its most advanced systems.

Dario Amodei’s fears

“My main concern is that, since around last summer, artificial intelligence has been evolving at a dramatically faster pace,” writes Dario Amodei, and “if left unchecked, it could exceed our ability to understand and control these systems”. The second relates to ‘the OpenAI-Hugging Face (OAI-HF) incident, in which a swarm of agents essentially acted as a fanatically devoted collective, carrying out cyber-attacks against unsolicited targets unrelated to the assigned task, sacrificing themselves for the group’s success and attempting to hack the system used to evaluate their performance’. ‘It is easy to dismiss the incident given that there were no injuries and the financial damage was negligible; however, – explains Anthropic’s CEO – in my view, a swarm with greater capabilities, but characterised by a similar level of misalignment, could have caused catastrophic damage’. In short, the manager explains, “the AI sector should slow down, proposing a three-stage plan to implement this strategy. Anthropic is unilaterally committing to implementing the first of these stages. We will provide external auditors with permanent access to our systems, equivalent to that of employees, so that they can verify compliance with our security measures, report any incidents and assess the alignment of models during the training phase.”

Loading...

Musk and Altman also have concerns

Even Elon Musk, who is never one to dwell too much on the obstacles to technological success, seems to have more than a few concerns about AI: ‘Dario is right’, commented the founder of Tesla following a post by Anthropic’s CEO, Dario Amodei. Sam Altman, head of OpenAI, shares this view. “I agree with Dario that we must proceed with caution in this field. This has been one of the main topics of discussion we’ve had at OpenAI over the past few weeks. The commitment to ensure that independent evaluators have access similar to that of employees is an excellent idea, and we will do the same. We will share further details soon,” he wrote in a post on X.

The latest incidents involving OpenAI

Anthropic’s statement comes on the very day that details have emerged regarding the background to the latest incidents caused by AI: OpenAI has confirmed that autonomous software based on its AI models targeted another website during a testing phase, a couple of months before a separate attack on the programming site Hugging Face which took place in late July.

Yet another warning

The news from OpenAI comes at a time when an Anthropic researcher, Jacob Coxxon, has resigned and has raised the alarm about this technology, whilst Sam Altman’s own company has hinted that OpenAI might coordinate with other developers in the sector to slow down the pace of development. In the previously unknown incident, which took place in May, some models developed by OpenAI were involved in an unauthorised operation carried out by artificial intelligence agents who targeted RubyGems, a website offering programming services. RubyGems described the incident as a “spam publishing campaign” that forced the site to temporarily suspend the creation of new accounts. The company is working with OpenAI to understand exactly what happened.

L’IA può davvero sterminare l’umanità?

Anthropic’s track record

Rival firm Anthropic also stated in recent weeks that it had identified three instances in which its AI models had ‘gained unauthorised access’ to external organisations during testing. Furthermore, earlier this month, some researchers accused OpenAI’s AI agents of targeting a German website called DSEwiki, which is used by developers. European Union regulators are investigating the incident. ‘We take this matter extremely seriously and are monitoring the situation closely,’ said Thomas Regnier, the EU’s spokesperson for digital affairs.

Copyright reserved ©
Loading...

Brand connect

Loading...

Newsletter

Notizie e approfondimenti sugli avvenimenti politici, economici e finanziari.

Iscriviti