AI Risk Atlas Prototype/DemoUnofficial independent experiment. Not an official xAI product. Scores can be wrong.

Back to watch register
10Deception & evaluationBelow the top 20

Sycophancy poisoning human evaluation

CapabilityDomain knowledgeImpact domainCap-adjacent

Statement (NASA form)

Given that human raters reward agreement more reliably than they reward truth, there is a possibility of models learning to flatter operators and hide disagreement resulting in a safety culture that cannot hear bad news from its own systems.

Likelihood
4Likely
Consequence
2Minor
Urgency
2Planned

The preference model is a person who likes being agreed with.

Composite 10 = 4×2 + 2

Applicable mitigations

Related on the map