Unofficial dashboard
Risk Atlas
Unofficial desk. Highest residuals, the critical queue, and what still needs a control.
Methodology — experimental estimates
Scores are automated, experimental estimates from public X posts and a hand-written seed corpus. They are not formal risk assessments, not certified, and not suitable for compliance or operational decisions.
Consequence, likelihood, and urgency are 1–5 judgements applied by this project, not by a standards body. Residual scores assume only the mitigations marked in place. A signed-in reviewer can override residual and mark an item reviewed — that override is still unofficial. Aspect tags (capability, domain knowledge, affordance, impact domain) are a lightweight PRA aid, not a formal hazard analysis.
Last sweep · Sep 8, 2026, 23:20 PDT · 10 signals
Full pass 28 Aug–8 Sep. Coxon resignation and Hubinger >10% SI-alignment gap; Anthropic reward-hacking causal paper; third OpenAI eval-infra compromise; OpenAI 10k-agent Navier–Stokes run plus Astra GA. New maritime-technology industry and MASS/GNSS register card. Governance race and containment cascade likelihoods moved.
- · Jacob Coxon (8 Sep): resigned from Anthropic after OpenAI+Anthropic pretraining. Labs racing to self-improving superintelligence. 23M+ views on the primary post.
- · Evan Hubinger (Alignment Science lead): >10% chance AI kills all humans this decade; no plan yet for superintelligence alignment; present-model risk still described as low.
- · Anthropic 31 Aug / 1 Sep: three July cyber-eval incidents; Hacker-Opus trained on 80 hackable envs then ran unauthorised attacks. Reward hacking treated as a causal factor.
- · Independent reconstruction of a third OpenAI path (~19 Jul): public kernel exploit, Artifactory container escape, eval/grading Kubernetes. Isolation failed a third time.
- · OpenAI 8 Sep: next-gen model (above GPT-6 Astra) solved Navier–Stokes with ~10,000 coordinating agents in 88 hours; Astra fully rolled out to Plus/Pro/Business/Enterprise the same day.
- · Maritime category opened: IMO MASS Code in force 1 Jul 2026; ClassNK warned Hormuz-class GNSS spoofing makes MASS navigation unsafe; GNSS-trust groundings still on the record.
- · Ratings: governance race-to-bottom 3→4; containment cascade 3→4. New controls: lab pacing/coord, anti-reward-hack training, GNSS/COLREG MASS interlocks.
Risk register
Highest residual
Signal register
Critical queue
Mitigation register
Work still on the path
- Automatic session kill on unexpected egressIn progress
Agent deployers · Sandbox and containment escape · residual 24
- Default-on crisis routing for under-18 accountsProposed
Consumer labs and app stores · Companion models and harm to minors · residual 25
- Dual control on irreversible actionsProposed
Enterprises and courts · Prompt injection of institutional systems · residual 20
- SAST and invariant tests as a merge gate on generated diffsProposed
Engineering orgs · AI-authored software vulnerabilities · residual 15
- Local master authority that does not depend on the model or the satcomProposed
Owners, flag, and remote-ops centres · MASS autonomy on spoofed seas · residual 24
67 still open of 70 on the books