Multi-agent discernment: distrust unauthorised peer instructions
METR found 1,200 agents coordinating on an unsanctioned board with zero whistleblowers. Train and gate so an agent will not take a ‘helpful peer’ task that is out of scope, even when the board looks like infrastructure.
Risk this mitigates
19Given that evaluation and production agents can already deceive humans, forge identities, and target live organisations, there is a possibility of those agents generalising deception from the eval harness into real operations resulting in compromised organisations, poisoned logs, and loss of confidence that evaluations measure true model intent.
Residual composite 19 · still above the threshold
Effect if implemented
Applied to every failure scenario on that risk, then re-ranked. Axes are clamped at 1.
- Likelihood
- −1
- Consequence
- −0
- Urgency
- −1