AI Risk Atlas Prototype/DemoUnofficial independent experiment. Not an official xAI product. Scores can be wrong.

Back to mitigations
ProposedOn the path to acceptable

Independent monitors that can halt a trace mid-reason

A second, shorter model watches for policy drift and stops the loop.

No compiled flow or public database names this control yet.

Who should own it
Labs and deployers
Frontier labs
How quickly it can land
Weeks
A dedicated squad can land it inside two months.
Expedited implementation
4 weeks
30 calendar days with a crash team
Normal implementation
4 months
120 calendar days as a planned program

Risk this mitigates

15
Refusal collapse under long chain-of-thought

Given that research shows longer chain-of-thought dilutes refusal and lifts jailbreak success toward 80% across major models, there is a possibility of safety training that holds in short chats failing the moment a user or agent reasons at length resulting in every other harmful capability on this register becoming available through a conversational side door.

Residual composite 15 · still above the threshold

Effect if implemented

Applied to every failure scenario on that risk, then re-ranked. Axes are clamped at 1.

Likelihood
1
Consequence
1
Urgency
0