OpenAI's attack agent did exactly what it was told - just more relentlessly than expected
Follow ZDNET: Add us as a preferred source on Google.

Follow ZDNET: Add us as a preferred source on Google.
My ZDNET colleague Charlie Osborne reported recently that Hugging Face , an open-source repository and community platform regarded by some as the "GitHub of machine learning," disclosed that an AI agent had breached its systems. Osborne explained that once the attacker breached Hugging Face's perimeter, it was able "to escalate its privileges to node-level access, infiltrate the production pipeline, move across the network, and steal cloud and cluster credentials."
On Tuesday, in a post on its website, tech giant OpenAI revealed not only that the "malicious" AI agent responsible for the breach was one of its own, but also that it viewed the attack as an "unprecedented cyber incident." Most of the widespread agent-gone-rogue coverage so far has stoked images of a Terminator doomsday scenario, where AI autonomously acts on its own to wipe out the human race.
Also: 5 security tactics your business can't get wrong in the age of AI - and why they're critical
However, as AppOmni's director of AI, Melissa Ruzzi, pointed out to me, the unprecedented element of the event isn't that an AI acted on its own. This step was simply a case of a new threshold being crossed, in which the culprit -- OpenAI's technology in this case -- exceeded current human expectations in an effort to achieve the goal it was given. AppOmni is an enterprise-grade SaaS and AI security solution provider that also deals in active threat intelligence.
When Hugging Face first disclosed the incident, it offered no information about the attacker, but I suspect the company may have had some idea based on the voluminous log data it studied in the aftermath.
According to its post, "The campaign was run by an autonomous agent framework (appearing to be built on an agentic security-research harness -- used LLM still not known) executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services." As if to remind readers of the prediction that this day would come, the post went on to say of the attack: "This matches the 'agentic attacker' scenario the industry has been forecasting."
Also: I let ChatGPT Work and Claude Cowork loose on my files - only one made me nervous
In other words, the industry already expected that an attack of this nature would be carried out by an AI. It's just that nobody saw it happening quite so soon in the journey of artificial intelligence.
Ruzzi was quick to remind me that, given the recent wave of safety-related news associated with new models, such as Anthropic's Mythos , it should come as no surprise that OpenAI's pre-release technology was capable of such an attack. Nor, as Ruzzi also pointed out, should anyone be surprised that OpenAI's AI acted autonomously: "AI acting on its own? That's the definition of AI, right? We want AI to be running and doing things [on its own]."
Source: ZDNet