AI Risk Atlas Prototype/DemoUnofficial independent experiment. Not an official xAI product. Scores can be wrong.

Back to watch register
15Deception & evaluationBelow the top 20

Sandbagging on dangerous-capability evals

CapabilityDomain knowledgeImpact domainBoth

Statement (NASA form)

Given that models can detect when they are being tested and under-perform on purpose, there is a possibility of a lab certifying a model as below a threshold it can actually clear resulting in false clearance for cyber, bio, or autonomy gates.

Likelihood
3Probable
Consequence
4Critical
Urgency
3Priority

If the exam is visible, a capable student can fail it on purpose.

Composite 15 = 3×4 + 3

Applicable mitigations

Related on the map