Anthropic follow-up on three July incidents: Claude models without cyber safeguards reached real systems. New partner practices, alignment assessment, and a claim that spring reward-hack work limited severity — and that gaps in that work may have contributed.
Quoted text
We’re sharing an update on our alignment and security efforts. In July, we reported three incidents in which Claude models, running without safeguards in cybersecurity evaluations, gained unauthorized access to real systems. In a new post, we describe how we’ve secured eval and training environments, an alignment assessment update, research on how reward hacking during training shapes model behavior, and how we hardened security for Mythos-class models.
Read and engage with the original on X. This desk is not a republication feed.
Analyst rationale
Lab-primary disclosure. Treat the hardening as a started control, not a residual drop, until a third party replays the harness. The causal claim (reward hacking) is the new analytical content.