
Distributed backdoors bypass local LLM safety monitors
Researchers have identified a critical vulnerability in runtime monitoring systems that protect multi-agent LLM deployments. By distributing malicious payloads across multiple agents, attackers can evade local safety checks that individually flag each component as benign. The work formalizes this as an observability boundary problem, proving that monitors operating on isolated message streams cannot detect coordinated harm that emerges only when fragments are reassembled. This finding exposes a fundamental gap in current deployment safeguards for tool-using systems, forcing a rethinking of how safety infrastructure scales to distributed architectures.68





















