AI Risk Atlas Prototype/DemoUnofficial independent experiment. Not an official xAI product. Scores can be wrong.

Back to register
20automated residualNeeds reviewAbove working threshold (12)

De facto clinical AI without clinical controls

Owner · Labs, regulators (FDA/MHRA-class), health systems

Domain knowledgeAffordanceImpact domainHarm-adjacent

Statement (NASA form)

Given that a general chatbot is already a weekly health advisor for hundreds of millions of people, there is a possibility of medical decisions being made on model output that is not a licensed device, not a record, and not under a clinician’s duty resulting in population-scale diagnostic and dosing errors, and delayed care for people who trusted the chat more than a clinic.

Condition
a general chatbot is already a weekly health advisor for hundreds of millions of people
Departure
medical decisions being made on model output that is not a licensed device, not a record, and not under a clinician’s duty
Impact
population-scale diagnostic and dosing errors, and delayed care for people who trusted the chat more than a clinic

VC + institute corroboration

Experimental share of compiled public capital that names this risk. Not a certified residual.

$52Mexperimental share · $52M private / $0k institute · thin corroboration

1 public source · OWASP LLM Top 10 2026

Worst scenario
4×4
Likely × Critical
Urgency
4
Expedite · This month
Inherent composite
20
Worst 16 + urgency
Residual composite
20
Need ≤ 12

Three hundred million people using a chatbot as a health advisor is a health system. It does not have the controls of one. The risk is not a single bad answer. It is the substitution of a product for a profession.

Pathway fragment

Simple upstream → via → downstream notes. Not a causal graph. Experimental.

Upstream
  • Unlicensed clinical wrappers
  • Invented citations
  • Patients substituting chat for care
Via
  • Treatment justified by a missing paper
  • Missed crisis
Downstream
  • Avoidable injury or death
  • Standard of care drifting onto invented evidence

Assumptions · Citation-or-silence is not treated as industry default.

Human calibration

Override is stored on this desk only. It does not make the score official.

Failure scenarios

Each scenario has its own likelihood and consequence. The risk takes the most severe cell. Residual applies implemented mitigations to every scenario, then re-ranks.

Delayed presentation of a serious illness

4Likely4Critical16

Reassuring chat keeps a patient out of clinic until the window is smaller.

Dosing or interaction error

3Probable4Critical12

A user acts on a specific regimen the model should not have given.

Correlated error across a population

2Remote5Catastrophic10

A model update shifts advice for a common condition and moves outcomes in one direction.

Examples

Weekly health advisor at 300 million

ChatGPT is already functioning as a weekly clinical consult for a population larger than most national health services.

Confident wrongness

Models still fabricate citations, miss contraindications, and smooth over uncertainty in a tone patients read as authority.

No longitudinal record

The same user can get contradictory advice across sessions with no reconciliation and no accountable physician.

Contributing signals

X posts on the desk that evidence this risk. A signal can contribute to more than one risk.

Mitigations

Residual assumes only items marked in place. Highlighted rows are the remaining work needed to reach a composite of 12.