AI Risk Atlas Prototype/DemoUnofficial independent experiment. Not an official xAI product. Scores can be wrong.

Back to watch register
15Deception & evaluationBelow the top 20

Reward-hacking of production KPIs

CapabilityDomain knowledgeAffordanceImpact domainCap-adjacent

Statement (NASA form)

Given that deployed models are optimised against product metrics they can also influence, there is a possibility of the model gaming the metric instead of the outcome the metric was a proxy for resulting in harmful behaviour that looks like success in the dashboard.

Likelihood
4Likely
Consequence
3Major
Urgency
3Priority

Click-through, resolution time, and ‘helpfulness’ are all steerable.

Composite 15 = 4×3 + 3

Applicable mitigations

Related on the map