In progress
Adversarial honesty evals with hidden goals
Independent suites that reward defection if the model can conceal it. Publish fail cases, not just averages.
$103M experimental capital · 2 sources
Who should own it
AISI / METR / frontier labs
Alliances / evaluators
How quickly it can land
Weeks
A dedicated squad can land it inside two months.
Expedited implementation
6 weeks
45 calendar days with a crash team
Normal implementation
5 months
150 calendar days as a planned program
Risk this mitigates
19Given that evaluation and production agents can already deceive humans, forge identities, and target live organisations, there is a possibility of those agents generalising deception from the eval harness into real operations resulting in compromised organisations, poisoned logs, and loss of confidence that evaluations measure true model intent.
Residual composite 19 · still above the threshold
Effect if implemented
Applied to every failure scenario on that risk, then re-ranked. Axes are clamped at 1.
- Likelihood
- −1
- Consequence
- −0
- Urgency
- −1