Article
AI News Agentic AI

Anthropic details near 200m exchange distillation campaigns from Chinese AI labs

The report attributes the largest effort to Alibaba, with a separate campaign it says routed requests from Chinese military networks targeting Claude's most capable model

by TechDefused Newsroom
The image depicts a hooded figure sitting at a computer with glowing red and orange code on screens in the background. The focus is on the back of the figure, emphasizing a mysterious and potentially illicit activity related to hacking. aiImage created using AI — Midjourney

Anthropic has released a report alleging persistent distillation attacks by China-based AI companies, saying it observed nearly 200 million exchanges tied to five separate campaigns.

What the campaigns were after

Anthropic, which builds the Claude family of large language models, said the campaigns targeted core capabilities including agentic capabilities, tool use, coding, data analysis and logical reasoning.

The company said the methods behind these campaigns had grown more sophisticated over time, writing, "Over the last several months, unauthorized labs have developed increasingly sophisticated methods to circumvent our defenses and harvest the capabilities of US frontier models."

The largest single campaign

Anthropic attributed the single largest effort to Alibaba, describing a campaign that produced 151 million exchanges between May and July 2026.

That activity was spread across roughly 3,500 accounts, and peaked at nearly three million exchanges in a single day.

A campaign tied to military networks

Anthropic separately described a Moonshot AI campaign that it says routed requests from Chinese military networks.

That campaign sent close to 300,000 requests to Claude over a 10-day span, using roughly 5,000 accounts, and primarily targeted Anthropic's Opus model, its most capable system.

How the attacks reportedly worked

The report says attackers focused on extracting a model's chain of thought, the step-by-step reasoning a model produces internally, in order to train smaller models through a technique called supervised fine-tuning.

Anthropic said it found prompts capable of tricking Claude into revealing that internal reasoning, despite the model's use of "summarized thinking" outputs, a feature intended to withhold the full underlying reasoning process from users.

Part of a longer pattern

The findings track a year-long escalation Anthropic has documented, moving from earlier allegations made in February to what the company describes as larger, more sustained campaigns through the middle of 2026.

That progression suggests the companies involved have continued refining their methods even as Anthropic has strengthened its defences, turning distillation into an ongoing point of friction between Anthropic and several major Chinese AI labs rather than an isolated incident.

by TechDefused Newsroom