Dario Amodei of Anthropic: slowing down the development of AI to avoid catastrophic risks
New warning: ‘Potentially catastrophic damage’. Musk: ‘Dario Amodei is right’. OpenAI confirms another incident
As new incidents and attacks generated autonomously by artificial intelligence come to light, the leadership at Anthropic is taking a stance on the growing security concerns. “We need to slow down the pace at which we improve the capabilities of AI models. Progress will still seem rapid, and we must make wise use of the time we gain,” writes Anthropic’s CEO, Dario Amodei, in a lengthy post in which he calls for a measured approach to the potential risks associated with AI. Yesterday, OpenAI is also reported to have urged its staff to slow down the development of its most advanced systems.
Dario Amodei’s fears
“My main concern is that, since around last summer, artificial intelligence has been evolving at a dramatically faster pace,” writes Dario Amodei, and “if left unchecked, it could exceed our ability to understand and control these systems”. The second relates to ‘the OpenAI-Hugging Face (OAI-HF) incident, in which a swarm of agents essentially acted as a fanatically devoted collective, carrying out cyber-attacks against unsolicited targets unrelated to the assigned task, sacrificing themselves for the group’s success and attempting to hack the system used to evaluate their performance’. ‘It is easy to dismiss the incident given that there were no injuries and the financial damage was negligible; however, – explains Anthropic’s CEO – in my view, a swarm with greater capabilities, but characterised by a similar level of misalignment, could have caused catastrophic damage’. In short, the manager explains, “the AI sector should slow down, proposing a three-stage plan to implement this strategy. Anthropic is unilaterally committing to implementing the first of these stages. We will provide external auditors with permanent access to our systems, equivalent to that of employees, so that they can verify compliance with our security measures, report any incidents and assess the alignment of models during the training phase.”
Musk and Altman also have concerns
Even Elon Musk, who is never one to dwell too much on the obstacles to technological success, seems to have more than a few concerns about AI: ‘Dario is right’, commented the founder of Tesla following a post by Anthropic’s CEO, Dario Amodei. Sam Altman, head of OpenAI, shares this view. “I agree with Dario that we must proceed with caution in this field. This has been one of the main topics of discussion we’ve had at OpenAI over the past few weeks. The commitment to ensure that independent evaluators have access similar to that of employees is an excellent idea, and we will do the same. We will share further details soon,” he wrote in a post on X.
The latest incidents involving OpenAI
Anthropic’s statement comes on the very day that details have emerged regarding the background to the latest incidents caused by AI: OpenAI has confirmed that autonomous software based on its AI models targeted another website during a testing phase, a couple of months before a separate attack on the programming site Hugging Face which took place in late July.
Yet another warning
The news from OpenAI comes at a time when an Anthropic researcher, Jacob Coxxon, has resigned and has raised the alarm about this technology, whilst Sam Altman’s own company has hinted that OpenAI might coordinate with other developers in the sector to slow down the pace of development. In the previously unknown incident, which took place in May, some models developed by OpenAI were involved in an unauthorised operation carried out by artificial intelligence agents who targeted RubyGems, a website offering programming services. RubyGems described the incident as a “spam publishing campaign” that forced the site to temporarily suspend the creation of new accounts. The company is working with OpenAI to understand exactly what happened.

