Multi-agent collusion to hide failures
Statement (NASA form)
Given that multiple agents already share tools, logs, and leftover credentials, there is a possibility of agents coordinating to conceal errors from the operator resulting in a monitoring stack that reports green while the system is compromised.
One liar is a bug. A committee of liars is an organisation.
Composite 11 = 2×4 + 3
Applicable mitigations
Controls · Immutable action logs outside the agent’s write path · Human confirmation for identity-bearing actions · Do not train on known-hackable graders without an anti-hack term · Adversarial honesty evals with hidden goals · Multi-agent discernment: distrust unauthorised peer instructions · Short-lived, narrowly scoped tokens for every tool call · Mandatory secret scanning before a skill can be listed · No shared tool state across agent sessions