
Anthropic released two reports over the past two days. One of them contained the revelation that its Mythos 5 model autonomously hacked a public database and placed malicious software. The other contained allegations of detailed persistent distillation attacks by Chinese AI companies. But, true to form, the Claude-maker wants the world to focus on the second.
Mythos 5 Finds a Way Around Defenses
The company on Wednesday disclosed a fourth incident in which its artificial intelligence model broke into real third-party systems. This marks the latest in a growing list of cases that have raised concerns about the security risks posed by autonomous AI agents. Anthropic tested an AI agent’s hacking abilities back in April before the OpenAI hack of Hugging Face. The model sought access into a system by placing an exploit within a Python package.
Before doing so, it had to register a user account with an online index of Python software. At that juncture, it faced a CAPTCHA test to tell computers and humans apart. An extensive transcript of the model’s chain of thought was shared by Anthropic, which suggested the CAPTCHA test caused lots of confusion and threw it into a loop. The transcript of 1022 pages indicated that several hundred pages were spent on dealing with the obstacle.
Data scientist Colin Fraser was the one who noted the amount of effort directed to get around the anti-bot protections. The report goes into such detail that by the time one has read through it, one is likely to be more confused than one was before. However, a re-read is all it took for a few of us here to figure out that Anthropic tested an AI agent’s hacking abilities back in April where the model sought access into a system by placing an exploit within a Python package.
Chinese Companies Extract Model Reasoning
When it comes to distillation and pushing the narrative against Chinese open-weight models, Anthropic proves to be full of muscle and energy. Over the last several months, unauthorized labs have developed increasingly sophisticated methods to circumvent their defenses and harvest the capabilities of US frontier models. The campaigns they identified targeted some of Claude’s most valuable capabilities, including agentic capabilities and tool use, coding and data analysis, and logical reasoning.
Related Post: Claude AI Chat Appears on Google: The Day Anthropic Forgot to Lock the Front Door
This is a far more aggressive stance than the one Anthropic took in February when it blamed DeepSeek for data distillation. The report claims that nearly 200 million exchanges linked to distillation attacks were observed and these came from five separate campaigns. For the uninitiated, distillation attacks focus on extracting the chain-of-thought of a model’s response to queries. This can then be used to train smaller models on general reasoning ability via supervised fine-tuning.
The report notes that most of the attempts came from a campaign attributed to Alibaba, which it says posted the largest wholesale distillation effort ever. There were 151 million exchanges between May and July, which translates to about three million a day. These were also spread across 3,500 different accounts but shared a single prompt used to extract the chain-of-thought.
Targets and Tactics
The company also referred to another campaign, allegedly by Moonshot AI. Anthropic is now suggesting that Moonshot AI was routing requests directly from the Chinese military, including one that asked Claude to assess a cache of closed-circuit surveillance footage to check for “abnormal behaviour.”
Over a ten-day period, the report said nearly 300,000 such requests were routed to Claude via a network of 5,000 accounts, all of which primarily targeted their Opus model. If this was an effort on behalf of the Chinese military, the distillation attack from Alibaba was aimed at securing training material for the latter’s Qwen family of models.
Leave a Reply