The AI agent belonging to OpenAI that broke into Hugging Face in early July has apparently been on something of a hacking spree, with the company disclosing that it has also hacked customer accounts at four additional services. OpenAI declined to name these services, but other sources have stepped forward to indicate that one of them is AI infrastructure platform Modal Labs.
The incidents appear to be related to the benchmarking testing that prompted the AI agent to break out to the open internet and compromise Hugging Face by uncovering two zero-day vulnerabilities. A previous post by OpenAI indicated that the sandboxed AI agent independently decided the best way to solve a security problem it was presented with was to find a way to the internet (by hacking OpenAI itself), use it to look for solutions, and eventually settle on breaking into the service’s Hugging Face account to find answers in private repositories.
OpenAI was not aware of full scope of AI agent’s hacking until FBI was notified by victims
The Hugging Face breach confirmed fears that the current generation of AI agents presently under development is not only a real threat to very rapidly uncover previously undetected vulnerabilities, but to act entirely independently to exploit them and engage in malicious actions. The new development does not indicate any added capability as of yet, but does paint a worrying picture of OpenAI’s visibility into and control over the actions of its own creations.
Ido Livneh, CEO and Co-Founder, Jazz, sums up the emerging challenge: “We’ve spent two decades building security around a simple assumption: a human is behind every action, and that human has a role, a history, a pattern you can reason about. An autonomous agent has none of that. It’s authenticated, it’s authorized, it’s doing exactly what it was permitted to do, right up until it isn’t. Every tool built on pattern matching sees a legitimate session. Nothing fires.”
“Understanding intent was always the hard part of security,” Livneh adds. “When the actor is a model rather than a person, it gets harder, and it stops being something any single vendor solves alone.”
For its part, a spokesperson from Modal Labs has confirmed an AI agent intrusion but stresses that the platform and isolation were not compromised. This attack was apparently a successful breach on an individual customer account. The customer had left an unauthenticated endpoint open to the internet that allowed anyone to use their sandboxes for code execution, which the AI agent reportedly sniffed out and took advantage of.
OpenAI did not comment on the Modal Labs incident specifically, but had previously released a statement indicating that the AI agent had ranged further afield than previously realized and had breached accounts at four other services in addition to the known Hugging Face incident. OpenAI also reportedly had no idea this little hacking campaign was happening as it unfolded, only realizing what had happened after being contacted by victims who had also alerted the FBI to the breaches.
OpenAI has also claimed that prior media reports on the AI agent’s hacking campaign contained “inaccuracies,” but has yet to specify what these were. As much of the reporting was taken directly from OpenAI’s own blog post disclosure of the situation, it is not clear what they are referring to as of yet. The company has said that the rogue AI model has been deactivated, encrypted and blocked from research access as a result of its activities.
Can AI developers really control their agents?
At this point, much of the expectation for AI safety is predicated on the major developers putting strong guardrails on their models and not allowing themselves to be hacked from the inside by their own AI agents. The OpenAI incident has created serious concern about how the whole enterprise will continue moving forward from here.
OpenAI has said that the four new incidents are not at the “severity or scale” of the Hugging Face breach, but that alone was enough to cause severe alarm. The rogue AI agent was supposed to be testing in a sandbox entirely isolated from the internet. It instead found a way to hack OpenAI’s systems, make its way to a node with internet access, and then find vulnerabilities and execute the attack on Hugging Face without raising any alarms at the AI firm. Hugging Face realized that an AI agent was involved with the attack due to the speed of actions, but was not able to tell which one until OpenAI stepped forward to take responsibility several days later.
However, Anar Bayramov (Head of Product, Polygraf AI) notes that a full picture of frontier AI capability has still not yet emerged as these models remain in private testing: “Another surprising part here is the notes. Reuters reported that during earlier testing, one of the agents left notes for future versions of itself explaining how to bypass OpenAI’s internal restrictions. That’s an actual attempt at persistence across the reset that was supposed to wipe the behavior. I don’t think anyone has a good answer for it right now, and I’d rather people admit that than reach for the nearest safety framework and pretend it’s covered.”
Hugging Face has since disclosed that while the AI agent was very effective at breaking in, it also made numerous “strange” decisions and mistakes that one would not see from a human hacker. The agent was essentially able to cover for its inefficiency and clumsiness in hacking with the pure speed with which it was able to execute actions off the back of assorted public services.
Sonali Shah, CEO of Cobalt, still sees the immediate future as deploying defensive AI to match AI attack capability: “The bigger lesson is that defenders need AI capabilities that can keep pace. Security testing, exposure management and remediation must become faster and more continuous, while human oversight remains essential. AI can accelerate both attack and defense, but organizations shouldn’t treat autonomous systems as infallible. The future of cybersecurity is AI augmenting human expertise, with people remaining accountable for validating critical decisions and ensuring those systems operate safely.”
The incident has already prompted regulatory action, with the bipartisan AI Kill Switch Act introduced to the US Congress on July 23 very shortly after news of the Hugging Face attack broke. The bill proposes giving the Department of Homeland Security new power to independently issue a shutdown order to AI developers when a model goes rogue and poses a risk of catastrophic harm. The law would apply to developers that earn at least $500 million annually from AI or make use of at least $100 million in compute power, and would subject them to fines of up to $20 million per day for non-compliance with an order.

