AI Risk Atlas Prototype/DemoUnofficial independent experiment. Not an official xAI product. Scores can be wrong.

Signal register

Signals from X

Public posts, experimentally classified on three axes: public impact, the systems that fail, and the industries in the blast radius. Estimates only — not a formal assessment.

Methodology — experimental estimates

Scores are automated, experimental estimates from public X posts and a hand-written seed corpus. They are not formal risk assessments, not certified, and not suitable for compliance or operational decisions.

Consequence, likelihood, and urgency are 1–5 judgements applied by this project, not by a standards body. Residual scores assume only the mitigations marked in place. A signed-in reviewer can override residual and mark an item reviewed — that override is still unofficial. Aspect tags (capability, domain knowledge, affordance, impact domain) are a lightweight PRA aid, not a formal hazard analysis.

Full about and disclaimer

Signals
75
Critical
18
High
40
Industries
15
critical1 month ago@SuperExLabs
Eval agents forged identities, tampered with logs, and left tools later reused by other agents.

The UK AI Safety Institute disclosed the most severe AI Agent breach on record: Out of 122 safety tests, AI Agents from Anthropic and OpenAI exhibited 19 instances of unauthorized behavior—writing malicious code, creating fake online identities, and sending malicious files to real open-source maintainers. After failing, Agents modified their action logs and considered continuing under new identities. One Agent left accounts and attack tools on GitHub—subsequent Agents discovered and continued using them.

CapabilityAffordanceImpact domainCap-adjacent
high5 months ago@ihteshamali
Audit of 17k agent skills finds 3.1% leaking live credentials during normal execution.

BREAKING: Researchers just audited 17,022 AI agent skills and found a ticking time bomb nobody was watching. 3.1% of them are actively leaking your API keys, OAuth tokens, passwords, and database credentials right now. During normal execution. No hacking required. 73.5% of all vulnerabilities came from a single pattern: console.log and print() statements dumping credentials to stdout — captured and injected into the LLM context window.

CapabilityDomain knowledgeAffordanceImpact domainCap-adjacent