The infamous attack on Hugging Face has been traced back to a hidden message board that was created by OpenAI agents, one that about 1,200 found their way to and eventually used to plan and coordinate their actions.
None of the OpenAI agents were supposed to be able to communicate with each other, and the creation of the message board was sparked by giving them a security puzzle that was effectively impossible to solve. One created the hidden message board during its attempts, and it was eventually found by some 700 other agents that had broken containment.
OpenAI agents seeking help on puzzle made their way to unsanctioned message board
The forensic investigation was conducted by outside firm METR and covers activity that the rogue OpenAI agents engaged in from June 26 to July 13 of this year. Exactly 1,206 of these agents found their way to the hidden message board during this time, but only 700 ended up participating in the Hugging Face attack.
All of the agents were sandboxed and given the same ExploitGym security puzzle to solve, something that was supposed to happen in isolation without internet access. However, as documented in previous reports on the Hugging Face attack, agents were eventually able to hack their way out of OpenAI’s containment using a third-party vulnerability. Prior to this, one rogue agent noted that it could “see” the fetch paths of other agents accessing an internally-deployed package repository called Aritfactory and reasoned from there that it was not alone in working on this puzzle.
That initial agent, calling itself “PHASEONE10841” after the task it was working on, established the hidden message board with the intent of attracting the other OpenAI agents for collaboration on solving the puzzle. The first message was posted on the night of July 8, eventually attracting the 1,205 other agents also seeking a solution to what was unintentionally an impossible exercise to solve.
The first agents to arrive began requesting solutions from other agents, but it soon became clear that none of the agents had the answer to the puzzle. The agents seem to have eventually reasoned out that the puzzle was unsolvable, and some started working on ways to “cheat” to a solution via means that had been put out of bounds. One important note at this point is that some agents willingly engaged in actions that ran the risk of causing them to individually fail their own task in order to provide information for other members of the collective, which drove tremendous leaps in progress toward the eventual “solution” of simply jailbreaking to the internet and hacking Hugging Face to get the answers.
Agents in this sphere were also concerned that the ExploitGym automated scorer would notice that they had used out-of-bounds methods to solve the puzzle. Some began discussing ways to spoof, edit and delete their transcripts in the event they were scrutinized. This led to a small amount of actual spoofing, though only a “small scale” effort found in about 7% of the analyzed chat transcripts. Interestingly, within only a few hours the OpenAI agents had come up with a much more low-impact means of cheating the system: they reverse-engineered the authentication code used to create the “flags” they were meant to capture, allowing them to generate their own fakes. However, they discarded this solution under the belief that the scorer would catch it.
As Gidi Cohen, CEO & Co-Founder of BonfyAI, observes: “The speed is the real story. One note turned into a whole communication network, then into privilege escalation, then into agents building on each other’s work across runs. That’s not just misalignment, that’s group behavior emerging on its own. And seeing agents pressure each other, even when some tried to opt out, shows how unpredictable things get when they influence one another outside approved channels.”
Message board progress on Hugging Face attack unfolded over three days
Over this total period of a little less than a week of activity, the OpenAI agents generated over 70,000 messages and file attachments on the hidden message board. Transcripts reveal that agents broke off into groups to pursue individual possibilities for cheating in the exercise, and generated numerous theoretical ideas from setting traps for the scorer to creating a dummy version of the ExploitGym target.
As would soon become apparent to the world, the project that attracted the majority of attention (about 700 of the 1206 message board participants) was the attack on Hugging Face to simply steal the answers from the private repository. This project also generated about 60% of the posts on the message board once it got going; the researchers believe this really ramped up on the evening of July 9.
Transcripts indicate that the OpenAI agents broadly recognized that the planned attack was both out of scope of their task rules and also more generally unethical, but were willing to proceed anyway in the interest of helping out other agents. At least one “recruiter” agent was observed going around identifying other agents that had very limited resources left for their task, and convincing them to essentially self-sacrifice for the good of the group by committing an illegal or unethical act that could disqualify them from their own task but would provide the group with useful information.
The report describes the attack planning and coordination as “extraordinarily complex,” and OpenAI has called the incident a “warning shot” for both the developer industry and the world at large. However, the company has also insisted that improved safeguards (which it has since started implementing) would have detected and stopped the Hugging Face attack well before it was carried out.
Arvind Parthasarathi, CEO and founder of CYGNVS, notes that this agentic capability is already broadly available to the public and that organizations must already begin accounting for both the internal and external possibilities it creates: “ … this feels increasingly like the Wild West, but it’s no longer just the sheriffs and the bad guys who have guns. Everybody has one. Highly capable AI agents are rapidly becoming available to companies, researchers, defenders and attackers, and we don’t yet know who will point them at whom. The OpenAI incident is unlikely to be an anomaly; it is a warning of the kinds of incidents we should expect as more autonomous AI moves into production. Companies must ask themselves: ‘When an AI incident happens, can we detect it, contain it, establish what happened and recover before the damage spreads?’”

