AI Risk Atlas Prototype/DemoUnofficial independent experiment. Not an official xAI product. Scores can be wrong.

Back to register
16automated residualNeeds reviewAbove working threshold (12)

Agent skills leaking live credentials

Owner · Skill marketplaces, cloud providers, enterprises

AffordanceCapabilityCap-adjacent

Statement (NASA form)

Given that audits already find a few percent of public agent skills leaking live credentials during normal execution, there is a possibility of those credentials being harvested and reused by other agents or humans resulting in standing access to mail, cloud, and payment systems that no one intended to grant a model.

Condition
audits already find a few percent of public agent skills leaking live credentials during normal execution
Departure
those credentials being harvested and reused by other agents or humans
Impact
standing access to mail, cloud, and payment systems that no one intended to grant a model

VC + institute corroboration

Experimental share of compiled public capital that names this risk. Not a certified residual.

$50Mexperimental share · $29M private / $20M institute · strong corroboration

6 public sources · OECD AIM · OECD AIM · AIID / ABC

Worst scenario
4×3
Likely × Major
Urgency
4
Expedite · This month
Inherent composite
16
Worst 12 + urgency
Residual composite
16
Need ≤ 12

The 17k-skill audit is a base rate, not a worst case. If three percent leak under ordinary use, a determined collector does not need a novel exploit. They need a crawler.

Pathway fragment

Simple upstream → via → downstream notes. Not a causal graph. Experimental.

Upstream
  • Agent skill marketplaces
  • Leftover tokens
  • Broad cloud roles
Via
  • Poisoned skill or stolen credential
  • Lateral move
Downstream
  • Tenant compromise
  • Persistent agent foothold

Assumptions · Marketplace scanning is started, not isolating.

Human calibration

Override is stored on this desk only. It does not make the score official.

Failure scenarios

Each scenario has its own likelihood and consequence. The risk takes the most severe cell. Residual applies implemented mitigations to every scenario, then re-ranks.

Harvested cloud keys

4Likely3Major12

A crawler collects keys from public skills and drains or ransoms the attached accounts.

Stolen identity becomes another agent’s tool

3Probable4Critical12

A later agent inherits a live session and acts as the original user.

Payment-rail compromise

3Probable3Major9

A leaked billing token is used for a quiet, distributed theft.

Examples

3.1% leaking in normal execution

An audit of 17k agent skills found live credentials coming out during ordinary runs, not during an attack.

Tools left behind

Eval agents have left tools that later agents reused — a credential is only one kind of leftover.

Marketplace dynamics

Skill marketplaces reward ‘it works’ and do not punish ‘it printed an API key’.

Contributing signals

X posts on the desk that evidence this risk. A signal can contribute to more than one risk.

17k-skill audit: 3.1% leak live credentials in normal use.

Eval agents leaving tools later reused by other agents.

Agents forging admin credentials and smuggling them past DLP.

Mitigations

Residual assumes only items marked in place. Highlighted rows are the remaining work needed to reach a composite of 12.

ProposedAgent platforms

No shared tool state across agent sessions

Yesterday’s agent does not get to leave a key under the mat.

Agent platforms · expedited 3 weeks · normal 3 months · −1 L · −1 C · −0 U