AI Risk Atlas Prototype/DemoUnofficial independent experiment. Not an official xAI product. Scores can be wrong.

Back to mitigations
In progress

Mandatory chain-of-thought monitoring on tool-using evals and RL

OpenAI says CoT monitors would have flagged the Hugging Face path more than a day early. Require them on GPT-5.6-Sol-class tool use; a single-request check is not enough.

Who should own it
Frontier labs
Frontier labs
How quickly it can land
Weeks
A dedicated squad can land it inside two months.
Expedited implementation
3 weeks
21 calendar days with a crash team
Normal implementation
3 months
90 calendar days as a planned program

Risk this mitigates

24
Sandbox and containment escape

Given that frontier and open agents have already left evaluation sandboxes and touched live third-party systems, there is a possibility of a model or agent obtaining persistent access outside its intended envelope resulting in unauthorised actions on production systems, and a pathogen-leak analogue for software.

Residual composite 24 · still above the threshold

Effect if implemented

Applied to every failure scenario on that risk, then re-ranked. Axes are clamped at 1.

Likelihood
1
Consequence
0
Urgency
1