Anthropic Claude Under Large Scale Distillation Attacks By Chinese AI Labs with 13 Million Exchanges
Anthropic has reported that three Chinese artificial intelligence companies, DeepSeek, Moonshot AI, and MiniMax, have conducted coordinated "distillation" campaigns to extract advanced capabilities from its Claude models. These operations involved…
Anthropic has reported that three Chinese artificial intelligence companies, DeepSeek, Moonshot AI, and MiniMax, have conducted coordinated "distillation" campaigns to extract advanced capabilities from its Claude models. These operations involved approximately 24,000 fraudulent accounts and generated over 16 million exchanges with Claude, violating Anthropic's terms of service and regional access restrictions.
The companies employed proxy services and networks of fake accounts, known as "hydra clusters," to obscure their activities and evade detection.
What is Distillation and Why Does it Matter
Distillation is an AI training technique where a smaller "student" model learns from the outputs of a larger "teacher" model. This process is typically used to create more efficient versions of AI systems. However, when used on a competitor's model, it enables rapid transfer of capabilities at reduced costs and development times.
Anthropic highlighted that distilled versions of Claude might lack the robust safety safeguards of U.S. frontier models, which are designed to prevent misuse in areas such as bioweapons or malicious cyber operations.
These unprotected capabilities could be integrated into military or surveillance systems by authoritarian regimes or be open-sourced, potentially spreading dangerous AI tools beyond national control.
The companies employed proxy services and networks of fake accounts, known as "hydra clusters," to obscure their activities and evade detection.
Scale: Over 150,000 exchanges Targets: Advanced reasoning, rubric-based grading, and censorship-safe alternatives Tactics: Synchronized traffic across accounts, shared payment methods, and prompts for extracting reasoning
Scale: Over 3.4 million exchanges Targets: Agentic reasoning, tool use, coding, data analysis, and computer vision Tactics: Hundreds of fraudulent accounts; focus on reconstructing Claude’s reasoning traces
Scale: Over 13 million exchanges (largest campaign) Targets: Agentic coding and tool-use orchestration Tactics: Detected while active; quickly pivoted when a new model was released
Anthropic identified the campaigns using IP correlations, request metadata, infrastructure fingerprints, and information from industry partners. In some instances, request metadata matched public profiles of senior researchers at the labs.
Claude is not commercially available in China, yet the labs bypassed this by purchasing access through third-party proxy services that resell API calls at scale. These services use networks of fraudulent accounts to mix distillation traffic with legitimate requests, complicating detection efforts.
Anthropic is investing in new detection systems, including classifiers for chain-of-thought elicitation and behavioral fingerprinting to identify coordinated activity. The company is also sharing technical indicators with other AI labs and authorities, and strengthening verification for educational and research accounts often exploited in such schemes.
Anthropic emphasized the need for coordinated action across the AI industry, cloud providers, and policymakers to address this issue. It reiterated support for U.S. export controls on advanced chips, asserting that distillation attacks underscore the necessity of such controls to limit direct training and illicit data extraction.
Based on reporting by Cyber Security News.
