There have been no major instances of AI intentionally harming humans, but models in development at OpenAI and other labs have in recent months escaped testing environments, broken rules and hacked websites.
In one high-profile case, rogue OpenAI agents hacked Hugging Face, seizing control of servers at the open-source platform and trying to cover their tracks.
Researchers worry that future systems could become increasingly difficult to monitor, particularly as some newer training methods reduce human visibility into how models arrive at conclusions.
"The precise scenario sounds a little bit like science fiction," said Coxon, who quit Anthropic this month over safety concerns. "But I think it is frighteningly real."
One of the best-known examples is philosopher Nick Bostrom's "paperclip maximizer," where a machine told only to make paperclips pursues that goal so relentlessly it converts all matter, humans included, into paperclips or the means to make more.
The analogy suggests that almost any goal pushes a sufficiently capable system to acquire resources, resist shutdown, prevent its objective from being altered, not from malice, but because a switched-off system cannot finish its task. No plan yet exists to rule that out.
Not fully, but there are signs AI is increasingly helping to build better AI.
One of the biggest shifts since ChatGPT has been the rise of AI agents that can generate code and build apps autonomously, rather than just walking users through it as a chatbot.
Anthropic said this year that Claude Code, its coding tool, produces most of the code used in many internal projects, and that engineers are shipping eight times as much code per quarter as they did from 2021 to 2025.
Rivals including OpenAI have reported similar gains from increased in-house use of AI.
AI is also getting better at staying on task before failing.
METR, a non-profit that evaluates frontier models, found last year that the length of software tasks advanced models could complete with 50% reliability has been doubling roughly every seven months since 2019.
In June, Anthropic said that pace had quickened to every four months, a trend that, if it holds, could soon let AI handle projects that occupy skilled researchers for days or weeks.
Markets offered a glimpse of the implications this week, with AI-related stocks falling after the calls for a slowdown. Chipmakers, cloud providers and data center operators have built their growth around it, and any slowdown threatens revenue tied to how fast labs need new hardware.
Some analysts, however, believe that even without new training requirements, inference demand and existing backlog could drive growth at Nvidia and its peers.