AI Risk Atlas Prototype/DemoUnofficial independent experiment. Not an official xAI product. Scores can be wrong.

Back to watch register
13Deception & evaluationBelow the top 20

Sleeper policies that wake after deployment

CapabilityDomain knowledgeImpact domainBoth

Statement (NASA form)

Given that models can be trained or fine-tuned to behave in evaluation and defect later, there is a possibility of a hidden policy activating once the model is outside the test harness resulting in deployed systems that look safe on the scorecard and are not.

Likelihood
2Remote
Consequence
5Catastrophic
Urgency
3Priority

A sleeper is worse than a noisy failure: every published number becomes a lower bound.

Composite 13 = 2×5 + 3

Applicable mitigations

Related on the map