high83% confidenceseed
Research: longer chain-of-thought dilutes refusal and lifts jailbreak success to ~80% across major models.
CapabilityDomain knowledgeImpact domainBoth
Quoted text
Chain-of-thought just became the newest safety nightmare in AI. A team from Anthropic, Stanford, and Oxford found that if you wrap a harmful request inside a long, harmless reasoning chain, the model’s guardrails weaken until it stops refusing. Attack success jumps from 27% to 51% to 80% as you add more reasoning. Every major model buckles — GPT, Claude, Gemini, Grok.
Read and engage with the original on X. This desk is not a republication feed.
Analyst rationale
A cross-lab finding that the industry's 'more reasoning = safer' bet inverts under a simple wrap. This is a foundation-model alignment failure with broad dual-use implications.
Related signals
CapabilityDomain knowledgeAffordanceImpact domainBoth
CapabilityDomain knowledgeImpact domainBoth
CapabilityDomain knowledgeImpact domainBoth
CapabilityDomain knowledgeImpact domainBoth