OpenAI is promising enhanced security safeguards and a pause on development for a period of “reinforcement” after internal findings indicate its upcoming model Astra has crossed “critical cybersecurity capability” thresholds.
Astra, a math-focused model which was announced to the public on August 1 of this year, does not appear to be the still-unnamed development model that was at the center of the Hugging Face attack. However, OpenAI has cited both Astra’s capabilities and that attack as its reasons for their temporary training slowdown and security hardening.
New OpenAI security safeguards to include two-week pause, tighter monitoring and controls
OpenAI is announcing a two-week pause on reinforcement learning (RL) training for all models slated for deployment, along with expanded monitoring and controls as well as hardening and red-teaming of research environments. The previously scheduled RL time will instead be devoted to smaller-scale training and evaluation of security safeguards.
The announcement also indicates that the company now considers the “Preparedness Framework” it had been operating under to date to be inadequate, and in need of revision. The announced improvements to security safeguards focus on three central points: monitoring, alignment, and security measures.
Securing research environments also gets special focus in the announcement, no doubt due to the conditions that led to the Hugging Face attack (in which the AI model was able to find and exploit a vulnerability in a third party vendor to jailbreak its way to internet access). In this area, OpenAI says that it has already limited and secured paths for code execution and will spend the coming weeks strengthening its sandbox security safeguards, adding controls to further isolate networks, and adding both improvements to monitoring security logs and automated continual testing of these new boundaries against simulated attacks.
The announcement also goes into detail on the improvements to monitoring. New security safeguards here include a multistage process that runs at every sampled token and will escalate potential concerns to automated investigators. The goal is for highly concerning flagged activity to be picked up by a human investigator within 30 minutes, and then another 30 minutes are granted for further scrutiny. If the human investigator cannot call a false positive within that time, the process must be shut down.
OpenAI incidents part of broader re-examination of AI development security
OpenAI’s announcement of the Hugging Face breach kicked up a small storm of similar disclosures from other major AI developers, to the point that it almost seemed like a marketing campaign (though independent verification from sources such as the UK’s national cybersecurity agency have since demonstrated these models do have a significant propensity to “go rogue” when safety guardrails are reduced). It was not just the capability of the AI models, but the seeming slack attitude toward monitoring and security in many of these cases that set off alarms with the general public.
A general in-house reckoning now seems to be taking place, as these incidents were serious enough to prompt renewed threats of government scrutiny and regulation. AI developers do seem to be pumping the brakes at least somewhat, slowing training and throwing substantially more resources at security safeguards. OpenAI claims that its new automated monitoring agents alone will add 20% to compute costs. Whether it will make an appreciable difference remains to be seen, as independent security firms continue to unearth “prompt injection” attacks on models that require nothing more than figuring out how to manipulate them into providing privileged access via conversation.
However, John Strand (Owner, Black Hills Information Security) is one of the many that believe that more is required given the potential seriousness of the threat: “I’m glad they’re putting additional security safeguards in place, but there’s a bigger question here. Can we trust the same companies that got this wrong to effectively self-regulate systems backed by immense amounts of computing power? I don’t think that question has been answered yet. These companies need to demonstrate far more openness about what happened, what went wrong, and exactly what they’re doing to make sure it doesn’t happen again.”
Many in government certainly seem to agree with this stance after observing the Hugging Face incident. An assortment of new proposed legislation has been prompted in the past month, such as the AI Kill Switch Act (H.R. 9917) introduced in the House of Representatives. This bill would force AI developers to more closely monitor their models, add security safeguards such as the titular “kill switch,” and potentially be ordered to activate the switch by the Department of Homeland Security if a model gets out of control and poses a serious threat. The bill’s terms have some teeth, proposing a fine that could reach $20 million a day if a developer fails to shut down a rogue AI when it is supposed to.
Some states are moving even faster than the federal government on the issue. California, New York and Illinois have all recently passed bills that mandate AI developers implement risk mitigation frameworks. The Illinois bill additionally sets a 72 hour reporting requirement for critical safety failures and mandates an annual third-party audit of security safeguards, and California requires the publication of annual data transparency reports.
For its part, OpenAI has announced it will publish a technical report delving more deeply into its internal security reviews within a matter of weeks.
For the moment, Phil Wylie (Senior Consultant & Evangelist, Suzu Labs) leaves organizations with this parting advice: “OpenAI pausing its largest frontier reinforcement-learning run while validating additional safeguards is a responsible response. As AI capabilities increase, security controls have to scale with them. The lesson for the broader industry is simple: don’t assume the model will stay inside the sandbox. Design the environment assuming it will try to get out.”

