Models accelerating their own R&D faster than review
Statement (NASA form)
Given that labs already use models to write kernels, evals, and training code, there is a possibility of capability jumps that no human safety case has time to cover resulting in a new system whose residual is last quarter’s number.
Anthropic’s own reports flag automated R&D. A scorecard that cannot keep up with the thing it scores is a gap.
Composite 18 = 3×5 + 3
Applicable mitigations
Controls · Immutable action logs outside the agent’s write path · Human confirmation for identity-bearing actions · Do not train on known-hackable graders without an anti-hack term · Adversarial honesty evals with hidden goals · Multi-agent discernment: distrust unauthorised peer instructions · Statutory independent safety function at designated labs · Protected channels for safety staff · Pacing agreement / temporary capability freeze among US labs