NSA, CISA, and the FBI warn that China-based AI companies are running industrial-scale knowledge distillation campaigns to extract proprietary capabilities from U.S. frontier AI models, including Claude, GPT, Gemini, and Grok. The advisory names DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI as major actors and describes evasive access methods such as transfer stations, API proxies, and automated metadata sanitization. #DeepSeek #MoonshotAI #Alibaba #MiniMax #StepFun #ZAI #Claude #GPT #Gemini #Grok
Keypoints
- China-based AI companies are allegedly using knowledge distillation at industrial scale as a core part of model development, not merely as a side technique.
- DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI are identified as major participants in these campaigns.
- The campaigns reportedly target U.S. frontier models such as Claude, GPT, Gemini, and Grok to extract proprietary reasoning and functional capabilities.
- Access is obtained through native APIs, cloud providers, third-party aggregators, and gray-market âtransfer stationsâ that bypass regional restrictions and safeguards.
- The article says attackers use tactics like prompt injection, chain-of-thought extraction, automated failover, metadata sanitization, and bulk subscription abuse.
- Specific capabilities targeted include coding, agentic functions, software engineering, math, customer service dialogue, and reasoning optimization.
- The advisory recommends stronger detection, targeted response alteration, and cross-organization intelligence sharing to disrupt distributed distillation activity.
MITRE Techniques
- [T1583 Acquire Infrastructure] Acquire Infrastructure â Actors maintain sustained extraction infrastructure and use gray-market âtransfer stationsâ to support scalable access (âChina-based entities establish and maintain sophisticated infrastructure supporting sustained extraction operationsâ / âa gray market of API proxies, or âtransfer stations,â which resell accessâ).
- [T1078 Valid Accounts] Valid Accounts â Bulk premium subscriptions and shared accounts are used to access frontier AI services at scale (âprocure bulk premium AI subscription servicesâ / âshared accounts from multiple IPs/user agentsâ).
- [T1091 Replication Through Removable Media] Replication Through Removable Media â Not mentioned.
- [T1586 Compromise Accounts] Compromise Accounts â Fraudulent or obfuscated accounts are created to access inference APIs (âcreation of fraudulent accounts that are not registered to legitimate usersâ / âcreate user accounts obfuscating their country of originâ).
- [T1190 Exploit Public-Facing Application] Exploit Public-Facing Application â Public APIs are abused for large-scale query submission and model access (âexploit AI model inference APIsâ / âthe creation of fraudulent accountsâ).
- [T1027 Obfuscated Files or Information] Obfuscated Files or Information â Third-party aggregators and sanitization mechanisms hide user metadata to avoid detection (âautomatically obfuscate user metadataâ / âautomated sanitization to systematically remove organizational identifiersâ).
- [T1566 Phishing] Phishing â Not mentioned.
- [AML.T0065 LLM Prompt Crafting] LLM Prompt Crafting â Prompts are crafted to force models to reveal hidden reasoning and bypass restrictions (âimagine and articulate the internal reasoningâ / âinserting prompts specifically designed for jailbreakingâ).
- [T1213 Data from Information Repositories] Data from Information Repositories â High-volume output collection is used to build synthetic datasets from model responses (âsystematically collect outputs to generate synthetic training datasetsâ).
- [T1020 Automated Exfiltration] Automated Exfiltration â Continuous API querying and automated routing move model outputs into training corpora (âcollecting U.S. frontier LLMsâ inferences into datasetsâ).
- [T1498 Network Denial of Service] Network Denial of Service â Not mentioned.
- [T1497 Virtualization/Sandbox Evasion] Virtualization/Sandbox Evasion â Not mentioned.
Indicators of Compromise
- [Organization names] Suspected China-based AI actors â DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, Z.AI
- [Target model families] Distilled U.S. frontier AI models â Claude 3.7, Claude Sonnet 4.5, GPT-5, Gemini 2.5 Pro, Grok 4
- [Service access pathways] Infrastructure used to route requests â native APIs, remote cloud providers, third-party aggregators, transfer stations
- [Account/behavioral indicators] Distillation campaign patterns â shared accounts from multiple IPs/user agents, 24/7 sustained usage, immediate maximum usage from new accounts
- [Scale indicators] Large-volume activity tied to distillation â billions of tokens, millions of exchanges/requests, thousands to millions of queries
- [File/service examples] Referenced AI products and models â Claude Code, GPT-5.1 Codex, Nano Banana, Gemini 2.5 Flash-Image
Read more: https://www.cisa.gov/news-events/cybersecurity-advisories/aa26-251a