China flag with code showing AI models

Chinese Actors Deploy LLM Gateways to Clone U.S. Frontier AI Models

A network of LLM gateways is enabling Chinese actors to bypass geographic restrictions, access frontier models, and potentially clone them.

According to Team Cymru, the entire ecosystem, consisting of over 80,000 servers, was designed to circumvent providers’ restrictions to enable malicious activity at scale.

The operators frequently use the open-source project Claude Relay Service (CRS 1.x) and its successor sub2api to access frontier models. With 26 commercial sponsors, Sub2api has been forked over 8,000 times on GitHub.

Over 80,000 LLM gateways deployed to bypass Frontier AI model restrictions

According to the researchers, the LLM gateways allow operators to create multiple accounts and route user requests through them to conceal the user’s identity. Over 457 different networks across various countries hosted the LLM gateways, widely distributing the routing infrastructure. One cluster hosted by a US-based VPS provider had over 4,000 addresses in China and Hong Kong, as well as 304 relays.

In eight days, the researchers identified 10,867 LLM gateways, with 9,456 running sub2api and 1,353 running Claude Relay Service. However, the number increased to 80,000 within a few days.

Claude Relay Service and sub2api enable attackers to combine multiple accounts or credentials into a common gateway. Consequently, the provider cannot determine the user’s identity as they can only determine the gateway’s IP address and the account credentials.

“While sub2api has more capabilities, both enable an actor to turn one set of AI accounts (subscription or API) into a shared gateway,” the team explained.

Consequently, the LLM gateways enable operators to share or resell credentials to third parties, bypass rate limits, usage metering, and regional restrictions, thereby enabling malicious activity and fraud at scale.

“A transfer station breaks the assumption every frontier-model control depends on: that the account making a request belongs to the party consuming the answer,” the team stated.

Additionally, the Chinese actors use outputs from American frontier AI models to train less powerful ones rather than developing them from scratch, a process known as model distillation.

“Rather than independently creating the research, data, and compute required to build a frontier model, an actor can query a stronger ‘teacher’ model at scale, collect its outputs, and use them to improve a cheaper ‘student’ model,” they wrote.

Chinese actors are cloning U.S. Frontier AI models at scale

The distillation campaign targets various models, including Anthropic, OpenAI, Google, and xAI. However, Anthropic was targeted more frequently.

In eight days, Team Cymru researchers observed attackers uploading 81GB of data to Anthropic using 17 LLM gateways and receiving 1.4GB in response, suggesting potential mass distillation activity. In late August 2026, they uploaded 14 terabytes and downloaded 7 TB within 8 days. However, the researchers could not view the submitted prompts to link the activity to model distillation.

In September 2026, federal authorities CISA, the NSA, and the FBI warned about Chinese AI companies conducting knowledge distillation attacks to train their own models. DeepSeek, a Chinese AI company, has been accused of extracting data from American frontier models to train its R1 and V3 models.

“China-based artificial intelligence (AI) companies are conducting systematic extraction of proprietary functionalities and capabilities of U.S. AI companies’ models through industrial-scale knowledge distillation campaigns that form the core—not merely a supplement—of their AI development strategy,” the joint advisory stated.

Other Chinese AI companies accused of model distillation attempts include Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI. While model distillation is a legitimate activity, the industrial scale and methods used by the Chinese actors have raised concerns about their intent.