Text settings Story text Size Small Standard Large Width * Standard Wide Links Standard Orange * Subscribers only Learn more Minimize to nav The United States has now named six Chinese AI firms accused of waging industrial-scale attacks distilling US frontier AI model capabilities and perhaps sparing billions in Chinese development costs. In a joint release Tuesday, the National Security Agency (NSA), Cybersecurity and Infrastructure Security Agency (CISA), and Federal Bureau of Investigation (FBI) alleged that DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI have been attacking US models since at least late 2024. The firms “likely” acted with “Chinese government awareness” when extracting capabilities from US models, including variants of Claude, GPT, Gemini, and Grok, agencies said. “China-based AI companies that conduct industrial-scale distillation against US AI models see significantly shorter AI development timelines and reduced financial expenditures in training a frontier model,” agencies said. All American AI firms must work with the government and US allies to end the alleged theft threatening the US lead in the AI race, the agencies said.
That will require coordinated action across the AI ecosystem to combat the “aggressive, malicious, and targeted distillation activities at an industrial scale that extract restricted proprietary functionalities and capabilities of US frontier AI models.” Attack methods include “exploiting AI model inference APIs” by bulk-buying fake accounts, agencies said. Not registered to legitimate users, these swarms of fraudulent accounts execute “highly coordinated queries featuring identical or similar prompt texts,” which range “from thousands to millions on similar topics.” Another common method is using prompt injection techniques to jailbreak models, including crafting “prompts forcing models to reveal their hidden [chain-of-thought] reasoning,” agencies said. For example, “DeepSeek employed prompts instructing models to imagine and articulate the internal reasoning behind completed responses and write it out step by step.” Fixes may frustrate AI users in US To encourage firms to work together, agencies recommended mitigations that would supposedly make it harder for Chinese firms to steal from US models. First, AI firms must improve detection of sophisticated campaigns that allegedly use tens of thousands of accounts relying on “a gray market of proxies” to evade geographical restrictions and “route distillation requests through multiple pathways to gain unauthorized access.” Flagging this activity should be somewhat easy, agencies suggested, since “campaigns span days to months with query volumes in the thousands to millions per domain, far exceeding legitimate research or development use cases.” Generally, they’ve recommended stepping up monitoring for “anomalous and malicious prompts, accounts, networks, and behaviors.” Because Chinese firms rely on “bulk procurement of the US AI companies’ premium subscriptions shared across teams of developers,” that effort should also include flagging accounts with suspicious subscription-to-usage ratios, as well as any new accounts immediately hitting maximum usage, agencies said.
Both indicate “bulk deployment with pre-engineered templates,” agencies said. US firms should also be strengthening “identity verification” of users and more closely tracking individuals using enterprise subscriptions (both of which potentially raise privacy red flags for legitimate users). Next, agencies asked firms to start dumbing down model responses when suspected distillation attacks are flagged. By “subtly” altering responses—such as by “presenting correct information with different reasoning,” adding stylistic inconsistencies, or reducing reasoning depth—firms can decrease the payoff for Chinese firms.
US firms could also secretly switch malicious accounts to an inferior model, and they should do so without providing any notice, agencies suggested. That particular mitigation step will likely be technically challenging. Agencies acknowledged, for example, that Chinese firms “employ aggressive, adaptive discovery to systematically identify valuable extractable data,” which they then collect to generate synthetic training datasets. Some firms can automatically detect when a smarter model is available and switch within 24 hours.
They also have automated quality assurance systems that detect when outputs are degraded and can otherwise differentiate ordinary “service issues from defensive data degradation,” agencies said. Also problematic: if US firms aren’t careful with targeting, any legitimate users perhaps caught up in the policing frenzy might be switched to a dumber model without receiving any alert. Or they could suddenly receive shorter responses or experience withheld capabilities, agencies acknowledged. Additionally, firms may possibly add “noise” to the output that restricts further queries.
Users will likely notice if outputs degrade, just like Chinese systems attacking models would. Last year, OpenAI quickly made changes to its automatic routing system after facing swift backlash when that system “consistently defaulted to less capable variants unless users explicitly added phrases like ‘think harder’ to their prompt,” Ars reported. Still, agencies think it’s best practice to “avoid informing China-based AI company users suspected of distillation campaigns of a switch to a downgraded model.” Acknowledging that such steps could frustrate users, agencies said that firms should try to “balance security with user experience” while accepting that some trade-offs, like “lower prediction precision and business usefulness,” may be inevitable to keep China from copying US capabilities. However, US firms should strive to ensure that “AI safety researchers and third-party evaluators” are “informed of model changes,” agencies suggested.
Finally, and seemingly most critical to the defense strategy long-term, agencies want AI firms and allied governments to share information to help leading firms track how distillation attacks evolve and avoid wasting time researching isolated anomalies. Cooperation is critical, the US thinks. If everyone cannot work together, then the US will face ongoing financial harm “through systematic extraction of proprietary functionality and capabilities, causing significant economic losses,” agencies warned. What did Chinese firms do?
AI firms have been warning about distillation attacks since last year. OpenAI accused DeepSeek of using data improperly, Google claimed attackers tried to clone Gemini, and Anthropic suggested that Alibaba should be criminally punished for allegedly launching the largest-ever cloning attack on Claude. Very quickly, the government got behind them, in April warning China that a crackdown was coming. The joint statement released on Tuesday, though, was the Trump administration’s “most detailed accusation yet,” NBC News noted.
In it, agencies claimed that stolen AI model capabilities “form the core—not merely a supplement”—of China’s AI development strategy. DeepSeek was accused of “extensive malicious distillation” on Claude, Gemini, GPT, and Grok models in efforts to “reduce its compute and research costs.” The Chinese firm allegedly took specialized training data and a range of capabilities, including agentic functions, assistant capabilities, writing optimization, question-and-answer optimization, and chain-of-thought reasoning. Moonshot AI took a similar approach, switching between models from leading US firms to distill fine-tuning techniques, reinforcement learning, software engineering, and math capabilities. Other firms, including Alibaba, MiniMax, StepFun, and Z.AI, seemed focused on copying particular models from Anthropic and OpenAI, agencies said.













Leave a Reply