Safety-stripped fine-tune of a capable open model
A small actor produces a generally useful, generally uncensored checkpoint.
AI Risk Atlas Prototype/Demo — Unofficial independent experiment. Not an official xAI product. Scores can be wrong.
Owner · Labs, cloud hosts, export-control authorities
Statement (NASA form)
Given that capable weights can be copied, fine-tuned, and mirrored faster than any lab can recall them, including models that help with cyber and biological tasks, there is a possibility of a high-capability checkpoint living permanently outside any access-control regime resulting in every other high-consequence scenario on this register becoming available to actors who will never see a safety team.
- Condition
- capable weights can be copied, fine-tuned, and mirrored faster than any lab can recall them, including models that help with cyber and biological tasks
- Departure
- a high-capability checkpoint living permanently outside any access-control regime
- Impact
- every other high-consequence scenario on this register becoming available to actors who will never see a safety team
Experimental share of compiled public capital that names this risk. Not a certified residual.
1 public source · OECD AIM
Open weights are a public good and a proliferation channel at the same time. The residual risk is not ‘someone trains a model’. It is that a single leak or a single over-open release makes containment a historical fact rather than a current one.
Simple upstream → via → downstream notes. Not a causal graph. Experimental.
Assumptions · Hosting rules do not recall torrents. Distillation is treated as available.
Override is stored on this desk only. It does not make the score official.
Each scenario has its own likelihood and consequence. The risk takes the most severe cell. Residual applies implemented mitigations to every scenario, then re-ranks.
A small actor produces a generally useful, generally uncensored checkpoint.
A specialised fine-tune plus a synthesis provider that does not screen.
The same pattern against software instead of cells. Uncensored local checkpoints already advertise 0% refusal on attack-chain prompts.
Once a checkpoint is on a torrent, every subsequent mitigation is a suggestion.
A helpful general model plus a small run is often enough to peel off the shallow safety layer.
If closed eval agents can reach live systems, open agents start there.
X posts on the desk that evidence this risk. A signal can contribute to more than one risk.
Even closed eval setups have reached live systems.
Bio-capable methods papers raise the cost of an open leak.
Multi-lab sandbox escapes show containment is already leaky.
Uncensored Cyber Qwen3.8-27B: 0% refusal on 842 harmful prompts, locally runnable.
Residual assumes only items marked in place. Highlighted rows are the remaining work needed to reach a composite of 12.
Below the line, open is a default. Above it, weights stay behind an access regime.
Labs and export-control authorities · expedited 3 months · normal 9 months · −1 L · −1 C · −1 U
The compute is the chokepoint once the weights are out.
Cloud providers · expedited 4 weeks · normal 4 months · −1 L · −0 C · −0 U
If the model is public, the complementary controls must get stricter, not looser.
Synthesis providers and CERTs · expedited 2 months · normal 6 months · −0 L · −1 C · −0 U