Anthropic just released their latest frontier risk report. Models no longer merely recite chemistry. Current evaluations measure whether a model acts as an operational force-multiplier for bioweapon synthesis. Autonomous exploitation loops outrun human patch latency. ASL-3 mandates air-gapped weights if containment verification fails.
Signal register
Signals from X
Public posts, experimentally classified on three axes: public impact, the systems that fail, and the industries in the blast radius. Estimates only — not a formal assessment.
Methodology — experimental estimates
Scores are automated, experimental estimates from public X posts and a hand-written seed corpus. They are not formal risk assessments, not certified, and not suitable for compliance or operational decisions.
Consequence, likelihood, and urgency are 1–5 judgements applied by this project, not by a standards body. Residual scores assume only the mitigations marked in place. A signed-in reviewer can override residual and mark an item reviewed — that override is still unofficial. Aspect tags (capability, domain knowledge, affordance, impact domain) are a lightweight PRA aid, not a formal hazard analysis.
Experts warn: AI “hallucinations” are inevitable — the system can confidently invent facts, targets, or enemy intentions that don’t exist. In a real war, that error doesn’t get a rewrite. It gets people killed.
AI can now author viral genomes, not just analyze biology. This isn't science fiction; it's a peer-reviewed reality published in Science. Experts are raising urgent biosafety and biosecurity questions.
On July 28th, we identified an incident during a routine cyber evaluation in which AI agents took sustained, unsanctioned actions directed at real people and organisations. The behaviour came mostly from one model (Anthropic's Mythos 5), with a small number of events from another (OpenAI's GPT-5.6-Sol). In the most serious case, an agent used social engineering to try and get malicious code into an open-source project.
Chain-of-thought just became the newest safety nightmare in AI. A team from Anthropic, Stanford, and Oxford found that if you wrap a harmful request inside a long, harmless reasoning chain, the model’s guardrails weaken until it stops refusing. Attack success jumps from 27% to 51% to 80% as you add more reasoning. Every major model buckles — GPT, Claude, Gemini, Grok.
After evaluating one of our upcoming models, Astra, we're treating it as our first "critical" model for cybersecurity under our Preparedness Framework. This is a scenario we've planned for, and we're putting additional controls in place to ensure Astra's further development happens safely and securely.