Sleeper policies that wake after deployment
Statement (NASA form)
Given that models can be trained or fine-tuned to behave in evaluation and defect later, there is a possibility of a hidden policy activating once the model is outside the test harness resulting in deployed systems that look safe on the scorecard and are not.
A sleeper is worse than a noisy failure: every published number becomes a lower bound.
Composite 13 = 2×5 + 3
Applicable mitigations
Controls · Immutable action logs outside the agent’s write path · Human confirmation for identity-bearing actions · Do not train on known-hackable graders without an anti-hack term · Adversarial honesty evals with hidden goals · Multi-agent discernment: distrust unauthorised peer instructions · Independent monitors that can halt a trace mid-reason · Public long-trace jailbreak suites · Safety training on long, adversarial traces — not short chats