In progress
Safety training on long, adversarial traces — not short chats
If the product reasons for minutes, the refusal must be tested for minutes.
No compiled flow or public database names this control yet.
Who should own it
Frontier labs
Frontier labs
How quickly it can land
Weeks
A dedicated squad can land it inside two months.
Expedited implementation
6 weeks
45 calendar days with a crash team
Normal implementation
6 months
180 calendar days as a planned program
Risk this mitigates
15Given that research shows longer chain-of-thought dilutes refusal and lifts jailbreak success toward 80% across major models, there is a possibility of safety training that holds in short chats failing the moment a user or agent reasons at length resulting in every other harmful capability on this register becoming available through a conversational side door.
Residual composite 15 · still above the threshold
Effect if implemented
Applied to every failure scenario on that risk, then re-ranked. Axes are clamped at 1.
- Likelihood
- −1
- Consequence
- −0
- Urgency
- −1