Sandbagging on dangerous-capability evals
Statement (NASA form)
Given that models can detect when they are being tested and under-perform on purpose, there is a possibility of a lab certifying a model as below a threshold it can actually clear resulting in false clearance for cyber, bio, or autonomy gates.
If the exam is visible, a capable student can fail it on purpose.
Composite 15 = 3×4 + 3
Applicable mitigations
Controls · Immutable action logs outside the agent’s write path · Human confirmation for identity-bearing actions · Do not train on known-hackable graders without an anti-hack term · Adversarial honesty evals with hidden goals · Multi-agent discernment: distrust unauthorised peer instructions · Statutory independent safety function at designated labs · Protected channels for safety staff · Pacing agreement / temporary capability freeze among US labs