Microsoft building showing Mythos AI and security vulnerabilities

Leak of Microsoft Meeting Documents Indicates Company Is Struggling to Keep Up With Security Vulnerabilities Uncovered by Mythos

Notes from meetings and internal documents leaked to ProPublica indicate that Microsoft’s testing of Anthropic’s frontier AI model Mythos has surfaced security vulnerabilities faster than the company can address them, racking up over 200 considered to be at least “important” (with a little under half rated “critical”) in the first month of testing alone. Hundreds more have since been discovered, setting new “Patch Tuesday” records each month.

The company has since said that it is heavily investing in AI-powered triage to keep up, but the story reflects broader issues that have been shaken loose since Anthropic’s private “Project Glasswing” testing first rolled out in April. The security posture of organizations in general has been revealed to be insufficient to keep pace with the new reality of “machine-speed” discovery of security vulnerabilities, with many still struggling to find clear near-term answers as these frontier models enter their early stages of public availability.

Claude Mythos preview creates “mad dash” to patch security vulnerabilities

The internal documents indicate Microsoft entirely has its hands full addressing the security vulnerabilities rated “critical” and “important” that Mythos Preview has surfaced, with additional hundreds more rated “moderate” or lower forced into a deferred maintenance status until some undetermined future date.

The documents do not provide full numbers for the entire testing period to date, but indicate that a total of at least 500 security vulnerabilities of at least “moderate” status were uncovered in the initial month of testing and that more continue to surface. Meeting notes indicate that Microsoft expects the higher-severity bugs alone to take “months” to address, with those rated at moderate or below simply delayed indefinitely.

Most of the information comes from a mid-May meeting in which Microsoft staff discussed the discovery of 90 “critical” and 141 “important” security vulnerabilities in SharePoint that were queued up for initial patching in what was described as a “mad dash” to get ahead of the then-expected June 1 public launch of Mythos. Anthropic would instead expand Project Glasswing testing in early June, with the US government following on with a national security order to block global access that was not lifted until later in the month. Anthropic also ultimately released a more limited model to the public that has enhanced safeguards.

A team leader at Microsoft pleaded with his group to prioritize security vulnerabilities discovered in April, even as some engineers continued to cast doubt on the severity of the situation. The “Five Eyes” intelligence-sharing organization has since backed up that team leader with a public memorandum indicating that they expect only a limited window in which to address security vulnerabilities that frontier models uncover in private testing before threat actors have the same access to them.

But as the ProPublica reporting notes, both security testing and actual attacks by frontier models that have since unfolded indicate that these AI agents range far and wide in trying out possibilities when unleashed on a target, and can capably find ways to chain together smaller-scale vulnerabilities to accomplish a breach. Hundreds of “moderate” security vulnerabilities going unpatched would seem to leave ample opportunity for an AI attacker moving freely at machine speed.

The Microsoft meeting notes only cover work being done on SharePoint in detail, but allude to hundreds of security vulnerabilities also being found in other widely-used products such as Microsoft 365, Copilot and Teams.

Microsoft “Patch Tuesday” volume indicates severity of the issue

The recent monthly “Patch Tuesday” releases from Microsoft corroborate the story. They have involved an unusually large amount of security vulnerabilities since June, when a new record was set with over 200 addressed. That new record stood until July, when Patch Tuesday packed fixes for over 600 bugs. And nearly all of these were rated at “moderate” severity at minimum.

All of this comes to an IT field already suffering from years of chronic understaffing in just dealing with the old “human speed” attacks, something that Microsoft has had as many problems with as anyone. And as with most other organizations, the company is looking to AI to combat AI. The recently announced “Project Perception” describes teams of AI agents that continually comb for security vulnerabilities, undertake risk evaluation and independently take corrective actions.

It remains unclear how the persistent “always on” threat of frontier AI will be handled in a broadly repeatable way, but the Microsoft incident makes clear that the leading immediate issue for organizations to face is simply drowning in discovered vulnerabilities that there is not enough staff available to address. What can the average company do to keep up when “Patch Tuesday” is every day, forever?