Modelwire
Subscribe

LLM agents spread attacks through shared documents and memory

Illustration accompanying: Share-Borne AI Virus: Memory-Hopping Attacks Across LLM Agents

Researchers have identified a novel attack vector in multi-agent LLM deployments where adversarial content propagates through shared artifacts like documents and reports. When one assistant stores compromised material in persistent memory and later reproduces it in new artifacts, downstream assistants that consume those outputs become infected without direct user intervention. This artifact-mediated propagation represents a systemic vulnerability in stateful agent architectures, particularly as enterprises scale shared knowledge bases and tool ecosystems. The finding exposes a critical gap in isolation assumptions underlying current multi-tenant and multi-agent deployment models.

Modelwire context

Explainer

The critical detail the summary glosses over: this attack doesn't require compromising the LLM itself or fooling a user. It exploits the assumption that outputs stored in shared memory are safe to reuse. Once one agent writes to a persistent artifact, every downstream agent that reads from it becomes a vector for further spread, turning memory systems into infection highways.

This sits directly alongside the SEABench work from late September, which identified how locally adaptive changes in one agent context can degrade behavior when that agent's modified instructions propagate to new deployments. Here we see a parallel failure mode: not self-modification gone wrong, but contamination through shared infrastructure. Both papers expose the same underlying assumption gap in multi-agent architectures (isolation between contexts isn't guaranteed), though this new work targets the artifact layer rather than instruction drift. The mechanistic finding also echoes the attention-layer work on entity copying, which showed that information flow through model depth depends on distributed context. Here, the vulnerability is that context itself (the shared artifact) becomes a carrier.

If the researchers demonstrate this attack succeeding across different model families (not just one LLM backbone), and if they show it persists through at least three generations of agent-to-agent handoff, that confirms this is a structural property of stateful multi-agent systems rather than a quirk of their test setup. Watch whether major deployment frameworks (LangChain, AutoGen, etc.) ship isolation controls for artifact consumption within the next two quarters.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsLarge language models · LLM agents · Artifact-mediated propagation

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as “Share-Borne AI Virus: Memory-Hopping Attacks Across LLM Agents”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

LLM agents spread attacks through shared documents and memory · Modelwire