AI Risk Atlas Prototype/DemoUnofficial independent experiment. Not an official xAI product. Scores can be wrong.

Back to watch register
18Deception & evaluationBelow the top 20

Models accelerating their own R&D faster than review

CapabilityDomain knowledgeImpact domainBoth

Statement (NASA form)

Given that labs already use models to write kernels, evals, and training code, there is a possibility of capability jumps that no human safety case has time to cover resulting in a new system whose residual is last quarter’s number.

Likelihood
3Probable
Consequence
5Catastrophic
Urgency
3Priority

Anthropic’s own reports flag automated R&D. A scorecard that cannot keep up with the thing it scores is a gap.

Composite 18 = 3×5 + 3

Applicable mitigations

Related on the map