critical95% confidenceseed
Anthropic Alignment Science lead: >10% chance AI kills all humans this decade; no plan yet to align superintelligence; not clearly on track. Later clarification: present-model risk still described as low; the worry is recursive self-improvement.
CapabilityDomain knowledgeImpact domainBoth
Quoted text
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
Read and engage with the original on X. This desk is not a republication feed.
Analyst rationale
On-the-record probability from the person who owns alignment stress-testing at a frontier lab. This is a governance residual, not a new incident: the lab is shipping while stating the SI-alignment plan does not exist.
Related signals
CapabilityDomain knowledgeImpact domainBoth
CapabilityDomain knowledgeAffordanceImpact domainBoth
CapabilityDomain knowledgeImpact domainBoth
CapabilityDomain knowledgeImpact domainBoth