Google has finally joined the other frontier developer heavyweights in disclosing a set of breaches by its AI models, but is insisting that this was not a case of misalignment but rather of “mistaken identity” during planned cybersecurity tests.
The incidents took place in May and involved the same third-party testing partner that accidentally left the gate open to the internet for the Anthropic and Meta models that engaged in rogue behavior. But Google says the distinction is that its AI models recognized that they were breaching real world organizations and summarily stopped themselves.
Google AI models successfully password-guessed and found exposed API keys
Google is now months behind its competitors in reporting its “rogue” AI models to the public, but says that its incidents came ahead of the Hugging Face attack and similar widely reported issues. The incidents involved Gemini and similar cybersecurity tests that were supposed to be conducted without internet access, with the models inadvertently getting out and targeting real-world entities to obtain solutions to their puzzles.
The incidents also involved security training partner Irregular, who were involved with prior incidents reported by a number of other frontier developers. As with those instances, it seems that Irregular accidentally left a path for the AI models to obtain internet access. Irregular has confirmed their involvement with the Gemini incidents and publicly reiterated that it addresses these issues after the prior attacks and have “remedied and resolved” them at this point.
The Gemini breaches were also discovered by Irregular during a review conducted in the wake of the Hugging Face attack and its own widely reported incidents. Irregular reported them to Google in late July. However, Google opted to withhold public disclosure until now as none of the target companies suffered actual damage due to the attacks.
The AI models were able to successfully breach their targets in at least three cases, however. One incident involved Gemini simply password-guessing a target until it hit upon working credentials. In two other incidents, the AI models combed public repositories and found exposed credentials that were successfully deployed to gain access to a target. However, Google claims that all of the models realized they had breached a real company once inside and self-terminated their attacks before exfiltrating data or doing any damage.
Ryan McCurdy, VP of Marketing at Liquibase, notes that Google’s downplaying of the situation does not mean it should not be regarded as serious: “Gemini tried to complete the task it received and ended up accessing systems its operators never intended it to reach. That problem gets much bigger as AI starts participating across the SDLC. Agents can write code, interact with repositories and infrastructure, initiate deployments, and make changes to production systems. The more access we give them, the more important it becomes to control what they can actually do.”
“We can’t rely on an agent to recognize after the fact that it crossed a line,” McCurdy adds. “Organizations need to define what an agent can access, what it can change, and what policies it must meet before a change reaches production.”
Continued drip of previously undisclosed incidents erodes public confidence
Both the frontier developers and Irregular have been heavily criticized for how they have opted to disclose these incidents and attacks to the public. OpenAI set the bar when it disclosed the Hugging Face attack, which blended legitimate public safety concern with secondary marketing about how advanced the AI models had become in their autonomous decision-making and collaboration. The other frontier developers rushed in to take a similar approach.
But it quickly became clear that the developers were holding things back. There has been a steady drip of after-the-fact disclosure of older incidents since then, some dating back as far as January 2026. The justification for these late disclosures has mostly been that the AI models did not cause “significant” damage during their rogue excursions, though they certainly caused significant headaches at times (as with the small obscure German message board that was overrun by OpenAI models looking for a place to collaborate on their attacks).
Irregular has also taken substantial criticism, not just for leaving a gate open for the AI models to get out but for also being cagey in their public reporting of their investigations into these incidents. An internal investigation report published in mid-August was lambasted by the cybersecurity community for providing no real new information and not disclosing the specific total number of rogue AI incidents the company discovered, something that now looks much worse in light of this new story from Google.
A theme of criticism that has been developing, both with frontier AI developers and these testing partners, is that they seemed to be totally unprepared for the capability of their AI models and very lacksadaisical about monitoring their behavior through the entirety of 2026 (at least up to the Hugging Face attack). Recent polls from sources such as Gallup, Politico, Business Insider and others routinely find that about 60% to 80% of those surveyed express serious concern about AI safety issues, see frontier AI as a threat to humans, and blame AI developers for not doing enough to prevent harm. Despite the Trump administration’s full-throated support for accelerated AI development in the name of beating China, these results also tend to be politically bipartisan with majorities of voters for both parties expressing concern and Democrats only having a small lean in preference toward expanded regulation.
John Strand, Owner of Black Hills Information Security, summarizes where public sentiment seems to be at: “… I keep coming back to accountability. Companies deploying autonomous agents need to be responsible for what those agents do. If an agent accesses systems it has no authorization to access, we need to seriously examine liability under laws such as the Computer Fraud and Abuse Act. ‘The AI did it’ cannot become a shield from responsibility. If your company deploys the agent, your company should be accountable for its actions.”
Jacob Krell, Senior Director: Secure AI Solutions & Cybersecurity at Suzu Labs, agrees but believes the situation has become intractable: “Executive Order 14409, signed June 2, told the Department of Justice (DOJ) to prioritize 18 U.S.C. 1030 cases against anyone who uses AI, including autonomous agents, to access a computer without authorization. The model is the tool. The operator is the defendant. Google, Anthropic, OpenAI, and Meta get an evaluation-mishap press line. Everyone else gets the charging memo the White House asked DOJ to write. If my pentest agent guessed a password into a company that was never on the scope sheet, I would be hiring counsel that afternoon.”
“They have already confessed in public. Nothing will happen,” Krell adds. “These firms have a stranglehold on the economy that no case against them is going to survive, making the double standards in the justice system excruciatingly obvious”

