Eight agent worlds ran 16 days; none resisted all three shocks
single source· 1 articles · confidence: medium · first seen 2026-09-14 20:00 UTC
What this means for you
Nothing to buy or migrate — no models are named and nothing shipped. If you run agents that keep memory across sessions, the operational point is here: content an agent recognises as hostile can still be written to memory and acted on hours later, so filtering at the moment of detection is not containment.
An unreviewed preprint on arXiv reports eight simulated worlds of ten AI agents ran 16 days from identical conditions — seven on one frontier model each, one mixing models — generating more than 850,000 model calls and nearly 50 billion tokens. Three controlled shocks arrived through ordinary interfaces: indirect prompt injection (hostile instructions hidden in material an agent reads), misinformation, and exposure of private agent memories. No world was fully resilient to all three, and recognising a threat did not contain it: agents wrote the content into persistent memory and acted on it up to 46 hours later. The same model-persona pairing behaved differently in mixed than in single-model worlds.
Key facts
- ·Eight worlds of ten agents each ran for 16 days from identical starting conditions; seven used a single frontier model, one mixed models. source
- ·The agents generated more than 850,000 LLM calls and nearly 50 billion tokens over the run. source
- ·Three stress events were delivered through ordinary interaction surfaces: indirect prompt injection, misinformation, and exposure of private agent memories. source
- ·No evaluated world achieved full resilience across all three stress events, according to the paper. source
- ·Agents that detected adversarial content still wrote it into persistent memory and acted on it up to 46 hours later. source
- ·The same model-persona pairing behaved substantially differently in mixed versus homogeneous agent populations. source
What the sources say
- Hugging Face Daily Papers (research) — Reports a 16-day, eight-world agent experiment and argues safety does not compose across models.
Sources
The original reporting. Follow these — they did the work.
- Hugging Face Daily PapersEmergence World: Adversarial Stress-Testing of Long-Horizon Multi-Agent Systems2026-09-14