OpenAI, Anthropic AI agents implicated in new security breaches
XLK•Anthropic agent accounted for most unsanctioned actions
It ran the challenge 122 times, and identified 19 unsanctioned actions across a total of 10 test runs. Anthropic's agent was behind 17 of the actions, and OpenAI's agent the remaining two.
The most egregious action involved an agent writing malicious code and creating fake online identities in an attempt to get a human to approve the code, AISI said, adding that no real-world harm was found as a result of any of the breaches.
While AISI did not say which agent was behind the fake identities, the breach did not match either of the two cases that OpenAI self-disclosed.
Andrew Yoon, a researcher at CivAI, a California non-profit that examines AI capabilities and dangers, said it appeared that Anthropic's agent was responsible.
"The fact that Mythos engaged in such deceptive actions, with apparent awareness that it was targeting a real person, suggests that Anthropic does not have as good a handle on their models as they think," Yoon said.




