In a dramatic twist in the AI world, Anthropic has blown the whistle on a series of distillation attacks from China-based AI companies, including Alibaba, Moonshot AI, and DeepSeek. As the competition heats up, these attacks have become more cunning, aiming to siphon off the crème de la crème of US frontier models' capabilities.
The report highlights how these unauthorised labs have been busy bees, developing clever methods to bypass defences and tap into the intellectual goldmine of models like Claude. The attacks zeroed in on extracting the chain of thought from model responses, a treasure trove for training smaller models on reasoning skills.
Anthropic’s report paints a vivid picture of nearly 200 million exchanges linked to these distillation campaigns. The culprits? Five distinct campaigns, with Alibaba leading the charge in what Anthropic describes as the largest distillation effort they've ever seen. Over a few months, Alibaba orchestrated a staggering 151 million exchanges, all with the goal of bolstering their Qwen models.
Moonshot AI, with its Kimi model, added a twist of intrigue by allegedly routing requests from the Chinese military. One eyebrow-raising request involved using Claude to analyse surveillance footage for abnormal behaviour. Talk about high stakes!
Anthropic’s revelations serve as a wake-up call for the AI industry, highlighting the need for robust defences against these sophisticated attacks. As the battle for AI supremacy rages on, the stakes have never been higher. Stay tuned, as this tech thriller unfolds!
Want to hear more? Join Mal & Matt on the Property AI Report Podcast each week!
Access from your preferred podcast provider by clicking here
Made with TRUST_AI - see the Charter: https://www.modelprop.co.uk/trust-ai
