China-Based Artificial Intelligence Companies Conducting Industrial-Scale Distillation Campaigns Against U.S. AI Companies

China-Based Artificial Intelligence Companies Conducting Industrial-Scale Distillation Campaigns Against U.S. AI Companies
NSA, CISA, and the FBI warn that China-based AI companies are running industrial-scale knowledge distillation campaigns to extract proprietary capabilities from U.S. frontier AI models, including Claude, GPT, Gemini, and Grok. The advisory names DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI as major actors and describes evasive access methods such as transfer stations, API proxies, and automated metadata sanitization. #DeepSeek #MoonshotAI #Alibaba #MiniMax #StepFun #ZAI #Claude #GPT #Gemini #Grok

Keypoints

  • China-based AI companies are allegedly using knowledge distillation at industrial scale as a core part of model development, not merely as a side technique.
  • DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI are identified as major participants in these campaigns.
  • The campaigns reportedly target U.S. frontier models such as Claude, GPT, Gemini, and Grok to extract proprietary reasoning and functional capabilities.
  • Access is obtained through native APIs, cloud providers, third-party aggregators, and gray-market “transfer stations” that bypass regional restrictions and safeguards.
  • The article says attackers use tactics like prompt injection, chain-of-thought extraction, automated failover, metadata sanitization, and bulk subscription abuse.
  • Specific capabilities targeted include coding, agentic functions, software engineering, math, customer service dialogue, and reasoning optimization.
  • The advisory recommends stronger detection, targeted response alteration, and cross-organization intelligence sharing to disrupt distributed distillation activity.

MITRE Techniques

  • [T1583 Acquire Infrastructure] Acquire Infrastructure – Actors maintain sustained extraction infrastructure and use gray-market “transfer stations” to support scalable access (‘China-based entities establish and maintain sophisticated infrastructure supporting sustained extraction operations’ / ‘a gray market of API proxies, or “transfer stations,” which resell access’).
  • [T1078 Valid Accounts] Valid Accounts – Bulk premium subscriptions and shared accounts are used to access frontier AI services at scale (‘procure bulk premium AI subscription services’ / ‘shared accounts from multiple IPs/user agents’).
  • [T1091 Replication Through Removable Media] Replication Through Removable Media – Not mentioned.
  • [T1586 Compromise Accounts] Compromise Accounts – Fraudulent or obfuscated accounts are created to access inference APIs (‘creation of fraudulent accounts that are not registered to legitimate users’ / ‘create user accounts obfuscating their country of origin’).
  • [T1190 Exploit Public-Facing Application] Exploit Public-Facing Application – Public APIs are abused for large-scale query submission and model access (‘exploit AI model inference APIs’ / ‘the creation of fraudulent accounts’).
  • [T1027 Obfuscated Files or Information] Obfuscated Files or Information – Third-party aggregators and sanitization mechanisms hide user metadata to avoid detection (‘automatically obfuscate user metadata’ / ‘automated sanitization to systematically remove organizational identifiers’).
  • [T1566 Phishing] Phishing – Not mentioned.
  • [AML.T0065 LLM Prompt Crafting] LLM Prompt Crafting – Prompts are crafted to force models to reveal hidden reasoning and bypass restrictions (‘imagine and articulate the internal reasoning’ / ‘inserting prompts specifically designed for jailbreaking’).
  • [T1213 Data from Information Repositories] Data from Information Repositories – High-volume output collection is used to build synthetic datasets from model responses (‘systematically collect outputs to generate synthetic training datasets’).
  • [T1020 Automated Exfiltration] Automated Exfiltration – Continuous API querying and automated routing move model outputs into training corpora (‘collecting U.S. frontier LLMs’ inferences into datasets’).
  • [T1498 Network Denial of Service] Network Denial of Service – Not mentioned.
  • [T1497 Virtualization/Sandbox Evasion] Virtualization/Sandbox Evasion – Not mentioned.

Indicators of Compromise

  • [Organization names] Suspected China-based AI actors – DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, Z.AI
  • [Target model families] Distilled U.S. frontier AI models – Claude 3.7, Claude Sonnet 4.5, GPT-5, Gemini 2.5 Pro, Grok 4
  • [Service access pathways] Infrastructure used to route requests – native APIs, remote cloud providers, third-party aggregators, transfer stations
  • [Account/behavioral indicators] Distillation campaign patterns – shared accounts from multiple IPs/user agents, 24/7 sustained usage, immediate maximum usage from new accounts
  • [Scale indicators] Large-volume activity tied to distillation – billions of tokens, millions of exchanges/requests, thousands to millions of queries
  • [File/service examples] Referenced AI products and models – Claude Code, GPT-5.1 Codex, Nano Banana, Gemini 2.5 Flash-Image


Read more: https://www.cisa.gov/news-events/cybersecurity-advisories/aa26-251a