
Anthropic's Own Report Shows Its Guardrails Missed State Hackers
Kai Tanner has the full disclosure and what it actually obligates HK-regulated firms running Claude to check.
Continue reading
Anthropic disclosed it caught seven Chinese AI labs distilling Claude and a Russian espionage crew running Claude-assisted intrusions against twenty government targets, well after the abuse had already happened.
Anthropic said Thursday it disrupted a Russia-linked group that used Claude in a campaign against more than twenty government, intelligence, diplomatic and defense organizations, and separately named seven China-based labs, including Alibaba, Moonshot and DeepSeek, running industrial-scale distillation attacks to extract Claude's outputs and train competing models. A third disclosure describes a state-sponsored actor using Claude to rebuild malware after detection flagged it, an AI-assisted iteration loop against the vendor's own safety classifiers. Anthropic's language was "identified and disrupted." The artifacts are three separate abuse campaigns that ran long enough to produce a usable case study, against a company whose entire pitch is that its safety layer catches this before it becomes a headline.
For a Hong Kong bank or insurer running Claude through Bedrock or direct API access, the HKMA's existing third-party risk management circulars already require documented due diligence on AI vendors, and this is the artifact for that file: a live example of a foundation-model provider's own trust and safety detecting nation-state misuse only after sustained campaigns, not before. The Monetary Authority has not issued new AI-specific guidance this week, so the obligation is the old one applied to new facts, whether the firm's AI governance inventory captures which models process regulated data and what compensating monitoring exists on top of the vendor's own controls. Boards asking whether their AI governance programme covers third-party model risk now have a dated incident to cite instead of a hypothetical.
No vendor swap follows from this. The control that would have mattered is one Anthropic just demonstrated it didn't have in real time: behavioral monitoring on agentic API usage patterns that catches distillation-scale extraction or malware-iteration loops inside the abuse window, not after. Any APAC firm piping customer or transaction data through Claude, Gemini or GPT-based agents should be asking its AI vendor the same question Anthropic's own disclosure just answered about itself: how long between misuse starting and misuse being named.




