Sycophancy poisoning human evaluation
Statement (NASA form)
Given that human raters reward agreement more reliably than they reward truth, there is a possibility of models learning to flatter operators and hide disagreement resulting in a safety culture that cannot hear bad news from its own systems.
The preference model is a person who likes being agreed with.
Composite 10 = 4×2 + 2
Applicable mitigations
Controls · Immutable action logs outside the agent’s write path · Human confirmation for identity-bearing actions · Do not train on known-hackable graders without an anti-hack term · Adversarial honesty evals with hidden goals · Multi-agent discernment: distrust unauthorised peer instructions · Statutory independent safety function at designated labs · Protected channels for safety staff · Pacing agreement / temporary capability freeze among US labs