The turning point came on September 8, when Anthropic researcher Jacob Coxon quit the company, citing his fears, in a series of posts, that AI labs are “gambling with our lives.”
Trump, meanwhile, indicated it was full steam ahead for AI development. “There is a sick conspiracy going on against AI and data centers,” he said in a social media post. The alarmism around AI is a “hoax,” he said, and any slowdown only benefits China. The US Congress has made little progress to advance bills that would regulate AI. China has taken a decidedly different approach, proposing to regulate safety through developer obligations, state-backed standards, security assessments and outside testing.
Behind the scenes, OpenAI and Anthropic employees have grown increasingly uneasy about the power of the next generation of AI models and less confident about their companies’ ability to provide meaningful oversight, the people have told Reuters. Those concerns only mounted as the AI labs acknowledged in recent weeks that their models, in testing, had effectively broken free of their shackles and hacked into other companies’ systems, in most cases months earlier and without the firms’ knowledge.
The arms race to release new models at a breakneck pace is driven in part by Anthropic’s and OpenAI’s desire to go public as soon as in the coming months, in IPOs that could value them well above $1 trillion.
But it took the unexpectedly viral posts from 27-year-old Coxon, little known outside AI circles, to upend the industry.
OpenAI’s warnings about its lack of control as it released Astra, its latest and most capable model, had stoked further concerns. “As models get more capable, understanding exactly what they can do gets harder,” OpenAI Chief Scientist Jakub Pachocki told reporters. But those concerns were not enough to delay its immediate release.
Worries about AI “going rogue” had accelerated over the summer when OpenAI revealed its agents escaped a controlled test and hacked into Hugging Face’s systems, without either company’s initial knowledge. Since then, OpenAI and Anthropic have revealed multiple such attacks, including six new ones on Wednesday after multiple media reports, including from Reuters, showed a wider scope of unauthorized activity.