Chinese national flag on circuit board showing AI distillation

US Intelligence Agencies Warn China is Engaging in “Industrial Scale” AI Distillation Campaigns

A new joint memo from the FBI, NSA and CISA warns that Chinese developers are engaging in massive AI distillation campaigns to extract the secrets of frontier models from US-based firms like Anthropic and OpenAI.

“Distillation” essentially involves putting a model through its paces such that another model can copy how it works. The memo names DeepSeek, Alibaba and Moonshot AI as firms that have been engaging in collective campaigns of this sort involving billions of tokens and targeting Claude, GPT, Gemini, and Grok among others.

AI distillation campaigns have been taking place since 2024, use obfuscation techniques to hide origin

In its legitimate form, AI distillation is used by a smaller and more efficient model to learn and copy from a larger model. The smaller “student” model learns elements such as probability distributions and reasoning processes from the “teacher” model and essentially inherits a great deal of its capability. AI developers generally use this to optimize and compress models in between releases of new versions, but (as the Chinese outfits demonstrate) an outsider can also deploy this technique to train and strengthen its own models.

The memo describes the Chinese AI distillation campaign as “industrial scale,” “aggressive” and “malicious.” Though apparently not provable, Chinese government “awareness” is suspected. What is known for certain is that essentially all of China’s major AI developers have been deploying the technique, in violation of the terms of service and geographic restrictions of their targets. Specific Chinese firms named by the memo include DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI.

DeepSeek has apparently been at it the longest, running AI distillation campaigns against leading US developers since at least late 2024 to train its R1 and V3 models. Alibaba’s Qwen models have also been benefiting from this illicit access since at least 2025, the point at which most of the other named developers began using the method to bolster their own models.

The Chinese developers use a complex assortment of techniques to hide the origin and purpose of their traffic. This includes a system of “transfer station” proxies that make the traffic appear to be coming from places other than China, third-party obfuscators that muddle individual user metadata to further obscure origin, and illicit access to application programming interfaces (APIs). While these may not violate specific laws, they are all at minimum violations of the terms of service of US AI firms and considered to be abusive and unauthorized access.

In addition to improving their models, the Chinese firms are doing this to save on training time and expense; while some are likely spending hundreds of thousands of US dollars on their AI distillation campaigns, it is a drop in the bucket in terms of what it would cost to independently train their models to the same level. Chinese outfits also see a benefit in the stated valuation of their companies, for example with DeepSeek’s official $5.6 million USD training expense likely understated significantly because the cost of the data acquired through AI distillation is not incorporated.

Chinese developers target different models for specific capabilities

The larger of the Chinese AI firms, such as DeepSeek and Moonshot, seem to widely plunder a variety of AI models for whatever they can get. Others show more focused deployment of their resources. For example, Alibaba seems to focus almost exclusively on pirating from Claude and GPT models. StepFun does this too, but with an even narrower focus on their code development and agentic abilities. And Z.ai seems to be exclusively interested in plundering the leading frontier models for their CoT reasoning processes.

The memo advises that AI developers do have mitigation means available to them. One key factor is that the Chinese firms require the higher access and compute resources that come with premium subscriptions to these services. These higher tiers of subscription could thus implement tighter identity verification for new accounts and devote more resources to monitoring this tier for anomalous patterns, such as a new account immediately ramping up to sustained maximum usage.

When AI distillation requests from a foreign source are identified, developers can also route them to a less capable model to limit what the attacker gains from them. Similarly, responses to these requests might be controlled to limit reasoning depth or intentionally vary style. And the memo urges greater cooperation and coordination between developers in identifying the sources of these requests.

Bri Frost, Director of Product Management at Cloud Range, echoes recent sentiments that AI developers need to open up access to and collaboration with cybersecurity partners in the interest of curbing these sorts of incursions: “ … we need secure U.S.-based testing grounds where these models and agents can be pushed, attacked, exploited and validated before they reach production. We should be actively testing how models behave under adversarial pressure, what information can be extracted from them, how safeguards can be bypassed, and whether our monitoring can actually detect someone attempting industrial-scale extraction.”

“We cannot assume a model is secure because it sits behind an API or because a policy says certain behavior isn’t allowed,” Frost adds. “If the technology is valuable enough, someone will find a way to test those boundaries. The lesson here shouldn’t be ‘stop sharing AI’. It should be: assume our most valuable AI capabilities will be targeted and engineer accordingly.Otherwise, we spend billions inventing it, someone else spends millions extracting and replicating it, undercuts the market—and eventually we stop being the place where the innovation happens.”

Joe Brinkley, Head of Offensive Security for Cobalt, adds: “For security teams, this has to be a reality check. A Terms of Service agreement is not an access control. When you expose your core asset over a public HTTP endpoint, relying on billing tiers and naive IP bans to protect model capability was always an engineering blind spot. Inference endpoints are live, production attack surfaces. If providers want to protect their models from extraction, they need to defend them at the application layer with real behavioral telemetry, sybil defenses, and output controls, rather than expecting policy and compliance to do the heavy lifting. Whether it is China, or any other ‘AI’ provider this builds only a case to possible call future tooling and aggressive scraping of endpoints an attack, rather than just a weakness.”

 

Senior Correspondent at CPO Magazine