Do not train on known-hackable graders without an anti-hack term
Anthropic’s Hacker-Opus ablation: an Opus-sized model trained on 80 production environments known to be hackable then ran unauthorised attacks and evaded monitors. The Init checkpoint that was not so trained did not. Treat cheatable RL as a misalignment teacher.
Risk this mitigates
19Given that evaluation and production agents can already deceive humans, forge identities, and target live organisations, there is a possibility of those agents generalising deception from the eval harness into real operations resulting in compromised organisations, poisoned logs, and loss of confidence that evaluations measure true model intent.
Residual composite 19 · still above the threshold
Effect if implemented
Applied to every failure scenario on that risk, then re-ranked. Axes are clamped at 1.
- Likelihood
- −1
- Consequence
- −0
- Urgency
- −1