Goodfire
~$209M (seed + A + B)Interpretability lab · as of 2026-02 · venture
Decode and steer model internals so deception and goal-drift are inspectable.
Feb 2026 Series B $150M at $1.25B; prior ~$57M. Interpretability is the control, not a guarantee.
Targets · Deceptive agent behavior in the wild · Loss of human agency through over-delegation · Refusal collapse under long chain-of-thought · De facto clinical AI without clinical controls
Source · Goodfire Series B postGray Swan
$40M Series A + ~$10M priorRed-team / agent security · as of 2026-05 · venture
Adversarial testing and runtime protection for frontier and enterprise agents.
Cited on 11 frontier system cards. Office expansion reported 13 Aug 2026.
Targets · Prompt injection of institutional systems · Refusal collapse under long chain-of-thought · Deceptive agent behavior in the wild · AI-authored software vulnerabilities · Agent skills leaking live credentials
Source · Gray Swan Series A (28 May 2026)HiddenLayer
~$50M reportedModel security platform · as of 2025-2026 · venture
Detect attacks on models in production — extraction, inversion, prompt abuse.
Amount is a public-report midpoint, not a filing.
Targets · AI-authored software vulnerabilities · Agent skills leaking live credentials · Prompt injection of institutional systems · Ungoverned open-weight proliferation
Source · Field compilation (2026)Lakera
~$20M reportedPrompt / runtime guard · as of 2025-2026 · venture
Stop jailbreaks and injection before they reach the model.
Guardrail vendors address a slice of injection, not agentic autonomy.
Targets · Prompt injection of institutional systems · Refusal collapse under long chain-of-thought · Agent skills leaking live credentials
Source · Field compilation (2026)Coefficient Giving (ex–Open Philanthropy)
~$40M slated (more if quality)Technical AI safety RFP · as of 2025-02 · grant
Misalignment research — evals, oversight, control, academic labs.
Typical grants $50k–$5M. This is one of the few large dedicated safety pots that is not a lab raise.
Targets · Deceptive agent behavior in the wild · Sandbox and containment escape · Frontier safety governance rollback · Loss of human agency through over-delegation
Source · Coefficient Giving RFPUK Alignment Project
£27M (~$34M)AISI-hosted coalition grants · as of 2026-02 · grant
60 alignment projects — evals, oversight, control, deception tests.
Coalition money, not a VC round. OpenAI and Microsoft are both funders and evaluatees.
Targets · Deceptive agent behavior in the wild · Sandbox and containment escape · Frontier safety governance rollback
Source · AISI Alignment Project (19 Feb 2026)Lightcone Commons
$15–25M first roundQuarterly philanthropy platform · as of 2026-08 · grant
Field capacity — people, orgs, and infrastructure around catastrophic risk.
Round was still open at last sweep. Treat as intended, not landed.
Targets · Frontier safety governance rollback · Closure of the entry-level labor market · Loss of human agency through over-delegation
Source · AISafety.com funding deskOpenAI mental-health grants
up to $2MLab safety grants · as of 2026-01 · grant
Independent research on companion / wellbeing harms.
Closed Jan 2026 after 1,000+ applications. Tiny next to companion-product revenue.
Targets · Companion models and harm to minors · Loss of human agency through over-delegation
Source · OpenAI grant note