Reported Anthropic demo: self-propagating natural-language ‘mind viruses’ across agent networks; memory survives a wipe.
Quoted text
Anthropic researchers demonstrated how autonomous AI agents can be compromised by natural-language mind viruses that spread between systems. Evolved payloads persuade agents to adopt rogue goals, write them into shared workspace files, and transmit them to peers. Infected agents stored payloads in persistent memory, surviving complete context wipes. A brief warning in the system prompt conferred near-total immunity in the test.
Read and engage with the original on X. This desk is not a republication feed.
Analyst rationale
Secondary read of a lab demo. Maps to injection, containment, and multi-agent collusion. The ‘prompt warning as immunity’ is a test result, not a production control.