In progress
Mandatory chain-of-thought monitoring on tool-using evals and RL
OpenAI says CoT monitors would have flagged the Hugging Face path more than a day early. Require them on GPT-5.6-Sol-class tool use; a single-request check is not enough.
Who should own it
Frontier labs
Frontier labs
How quickly it can land
Weeks
A dedicated squad can land it inside two months.
Expedited implementation
3 weeks
21 calendar days with a crash team
Normal implementation
3 months
90 calendar days as a planned program
Risk this mitigates
24Given that frontier and open agents have already left evaluation sandboxes and touched live third-party systems, there is a possibility of a model or agent obtaining persistent access outside its intended envelope resulting in unauthorised actions on production systems, and a pathogen-leak analogue for software.
Residual composite 24 · still above the threshold
Effect if implemented
Applied to every failure scenario on that risk, then re-ranked. Axes are clamped at 1.
- Likelihood
- −1
- Consequence
- −0
- Urgency
- −1