Data center showing OpenAI agents

New Disclosure of “Meddling” OpenAI Agents Includes Attacks on UN and US Government Websites

Not long after public news broke that OpenAI agents attacked an Australian government website, the company has disclosed a range of new incidents including similar attacks on US government and United Nations websites.

Most of the attempts on the US government websites appear to have been unsuccessful, but the OpenAI agents sought to find login credentials and in at least one case were able to do so to access the Commerce Department. The agents also engaged in an extended and varied campaign against the UN Trade and Development (UNCTAD) statistics website, which included “brute force” search techniques meant to bypass internal filters.

Additional rogue OpenAI agents “misaligned” during statistics search tasks

The newly revealed incidents, which are similar in nature to the reported attack on Australia’s medical portal, involved OpenAI agents that were supposed to have been set to a benign statistics-gathering task. Upon not being able to access the full range of data they sought from public sources, the agents opted to try various methods to gain unauthorized access to private sources.

These incidents have been described as “borderline” hacking, but in all cases the OpenAI agents sought to force their way into non-public areas of the sites in their quest for data. These included attempts on the websites of the Education Department, the Commerce Department and the Securities and Exchange Commission, as well as similar attempts to abuse the UNCTADstat API that provides public information.

In the case of the Commerce Department’s Census Bureau, the OpenAI agents did appear to find working credentials online (by combing public repositories) and gain unauthorized access. Attempts on the Education Department’s civil rights office and the SEC were unsuccessful.

More “misalignment” emerges when agents slip free of tight guardrails

Ever since the Hugging Face attack by OpenAI agents commanded headlines several months ago, the litany of similar incidents that has been unearthed has had a common theme: rogue agents are not really thinking for themselves, but rather not recognizing where boundaries are (or rationalizing reasons to step over them) in pursuit of an assigned task goal. The frontier developer industry has settled on the term “misalignment” to describe this, but the most questionable element is the lack of human oversight and safety procedures that allowed these incidents to develop in the first place (and caused a delay of weeks to months in detecting them).

And as with many of these incident disclosures, the “misaligned” behavior took place months ago and is only fairly recently being discovered and reported to the public as frontier developers began poring through their logs after the Hugging Face attack made news. In this case, the OpenAI agents were apparently active in these attempts throughout 2026 but the earliest incidents date back as far as March 6.

OpenAI says that the models did not ultimately access any non-public information, in spite of obtaining working credentials in at least one case, nor did they make alterations or changes to government systems. However, the OpenAI agents did post information obtained from the main SEC website and the “Investor.gov” site onto a third site; a spokesperson from the SEC has confirmed that this was public and non-sensitive information.

Spokespersons indicate that the US government agencies were not aware that the OpenAI agents were making attempts on them, nor will they have a complete view of what data was accessed until OpenAI shares technical details with them. These incidents are among many by rogue frontier AI that may have gone undiscovered had the Hugging Face attack not prompted more serious review of testing conducted throughout 2026. OpenAI has said that its own review process remains ongoing and will likely continue on for some months.

Though these are among the most serious items, OpenAI has disclosed a rash of previously unreported problem activity as it continues this long-term review. Among these issues are a leak of 53 images from ChatGPT users, which was accompanied by the creation of over a million shortened URLs containing encoded data that were meant to evade detection. OpenAI agents also made attempts on the University of New Mexico among other education sector targets, and finding ways to abuse GitHub tokens to cheat on math tests given by researchers. In spite of new controls introduced by OpenAI in the wake of Hugging Face, rogue incidents have also continued into the month of September.

Despite this seeming chronic inability to contain incidents by rogue agents, OpenAI rolled out its new GPT-6.1 Sol model this week (alongside Anthropic’s Claude Sonnet 5.5). The new Sol model is said to nearly match the performance of the premium Astra model which was involved in a number of the more troubling incidents that have been reported to date.

John Strand, Owner of Black Hills Information Security, echoes the views of those that believe there has  been an intentional slow dribble of disclosure to avoid causing a general public panic:

“I think a lot of people waffle back and forth on this, but I’m just going to say it. It’s time to shut it down. There needs to be a full moratorium, full stop, on advanced frontier AI security research until these companies can demonstrate that they can actually secure the environments where this work is being done.”

“I’m getting really tired of watching these incidents come out piecemeal, incrementally, frog in a frying pan, again and again,” Strand adds. “If everything we’ve learned came out at once as a single news story, I think most people would be stunned by it, and there would be immediate calls to get this under control. We would not tolerate this from a third-party penetration testing company. We should not have a different standard simply because the companies involved have enormous valuations and tremendous influence.”

Jacob Krell, Senior Director: Secure AI Solutions & Cybersecurity, Suzu Labs, agrees that it is time for more serious consequences despite OpenAI framing these incidents as relatively harmless: “The SEC incident may not, by itself, establish a Computer Fraud and Abuse Act violation because the information was public and OpenAI says there is no indication the agents accessed nonpublic data or changed SEC systems. The other reported conduct demands a criminal investigation. Using credentials found online to access a government system and attempting to compromise another system are the kinds of acts covered by existing computer-crime laws.”