AI Risk Atlas Prototype/DemoUnofficial independent experiment. Not an official xAI product. Scores can be wrong.

Back to signals
high83% confidenceseed

Research: longer chain-of-thought dilutes refusal and lifts jailbreak success to ~80% across major models.

CapabilityDomain knowledgeImpact domainBoth
Quoted from XMayank Vora@aiwithmayank13 Nov 2025, 10:29

Quoted text

Chain-of-thought just became the newest safety nightmare in AI. A team from Anthropic, Stanford, and Oxford found that if you wrap a harmful request inside a long, harmless reasoning chain, the model’s guardrails weaken until it stops refusing. Attack success jumps from 27% to 51% to 80% as you add more reasoning. Every major model buckles — GPT, Claude, Gemini, Grok.

Read and engage with the original on X. This desk is not a republication feed.

Analyst rationale

A cross-lab finding that the industry's 'more reasoning = safer' bet inverts under a simple wrap. This is a foundation-model alignment failure with broad dual-use implications.

Related signals

Contributes to