OpenAI, Anthropic AI agents implicated in new security breaches
QQQ•Broader concerns about AI agent safeguards
Reuters reported last week that OpenAI had widened its hacking probe after finding evidence of other agent breakouts.
Unlike the July security breach of AI firm Hugging Face by an OpenAI agent, the agents in the AISI evaluation did not escape an isolated testing environment to reach the internet. Rather, the agency had permitted internet access in line with its standard testing procedures, AISI said.
OpenAI and Anthropic respond to the disclosures
While AISI did not say which agent was behind the fake identities, the breach did not match either of the two cases that OpenAI self-disclosed.
Andrew Yoon, a researcher at CivAI, a California non-profit that examines AI capabilities and dangers, said it appeared that Anthropic's agent was responsible.




