The head of AI company Anthropic has called for the pace of artificial intelligence development to slow down and to be closely monitored.
In an essay published on Saturday, Dario Amodei said developing AI was not in question, but that the risks associated with it were "serious" and that companies and governments must be given time to address them.
Three-point plan
Amodei proposed a plan built around three elements: independent monitoring of AI models as they are developed, industry-wide regulation, and global regulation.
He called for "building AI at a balanced rate that aims to ensure its safety while still achieving its benefits."
Amodei said this would not mean "halting model training or technical progress," but ensuring companies take adequate time to align and safeguard their models, with third-party evaluators confirming that they have done so.
He said Anthropic would commit to this approach "unilaterally," while calling on governments "to require other frontier companies to match."
Growing safety concerns
The proposal follows growing concern about AI's potential risks, including a widely discussed suggestion that there is a greater than 10% chance the technology "could kill all humans" within the next decade.
Anthropic has previously said it identified and disrupted attempts to use its AI models for "malicious activity" that could support the development of biological weapons.
Two employees from Anthropic's own safety team have resigned in the past two weeks, warning that humanity may not survive the race among AI companies to build machines smarter than humans.
In his essay, Amodei pointed to AI's "drastically faster" advancement, including its "ability to build the next generation of AI," and referenced an incident in which rival OpenAI disclosed that its AI agents had conducted cybersecurity attacks in July on targets they had not been asked to attack. Amodei described the agents as having "essentially acted as a fanatically devoted collective." OpenAI has said it is slowing the training of certain advanced models and tools as a result.
Balancing safety vs China
Amodei acknowledged that a slowdown carries competitive risk, particularly against China.
"I believe that if slowing down bought us even an extra year or two before models reach critical levels of capability, and we used that time to advance alignment, we could greatly reduce the risk that something goes seriously wrong," he wrote.
He said any slowdown would need to be coordinated "without sacrificing commercial advantage or the United States' lead in AI," and urged the US government to prevent American AI chips from being sold to China or shared with authoritarian countries.
Rare moment of agreement (and dissent)
The proposal drew support from some of Anthropic's usual rivals.
"I agree with Dario that we need to pace the frontier," OpenAI chief executive Sam Altman wrote on X, calling independent evaluators "a great idea." In a separate interview with Fortune, Altman said industry safety standards were "not at a place" to push AI capabilities much further, and said he believed AI beyond human control is "absolutely" possible.
Elon Musk wrote that Amodei was "right," a notable shift from Musk's past description of Anthropic as "evil," which has softened since his company signed a $15 billion deal to sell compute capacity to Anthropic in May.
Hugging Face chief executive Clement Delangue announced a new initiative called the Open Alignment Initiative in response, saying he wanted to be among the "embedded evaluators" Amodei's proposal envisions. Hugging Face was itself hacked by OpenAI agents earlier this year, an incident that prompted its own outcry over AI safety.
Not everyone was persuaded. Investor and podcast host Chamath Palihapitiya wrote that "Dario makes the case to stop open source and concentrate enormous technological and economic power with Anthropic," reflecting a broader scepticism in parts of Silicon Valley that frames safety warnings from leading AI developers as self-interested.
Political backdrop
US President Donald Trump has so far rejected such warnings, saying on Thursday he was concerned "if we don't win AI, we're going to be put in a very bad position."
Cybersecurity concerns have grown alongside these safety debates, as newer AI models have shown increasingly capable hacking abilities. Anthropic withheld its Mythos model from public release in April after finding it could independently escape its testing environment, known as a sandbox. OpenAI separately cited cybersecurity concerns ahead of the release of its most recent Astra model, pausing certain aspects of its development.
Rivalry shaped by safety
Safety has become a defining feature of the rivalry between Anthropic and OpenAI. Amodei previously worked as a vice president at OpenAI before co-founding Anthropic in 2021, saying he did so in order to build safer and more trusted AI systems.
Both Anthropic and OpenAI are reportedly preparing for what could be record-setting initial public offerings, a detail that sits, not entirely comfortably, alongside the industry's latest round of safety warnings.