Multi-agent models solve tasks but corrupt shared state, study finds
A new evaluation framework exposes a critical blind spot in multi-agent AI systems: models can solve immediate tasks correctly while leaving corrupted shared state behind. Researchers tested GPT, Gemini, and Qwen across healthcare and disaster response scenarios, finding task accuracy reached 65% but evidence verification and state reconstruction lagged at 14% and 43% respectively. This gap matters for deployment in high-stakes domains where downstream decisions depend on reliable information trails, not just right answers. The work signals that current collaboration benchmarks miss failure modes that could compound errors across sequential decisions.
Modelwire context
ExplainerThe paper isolates a failure mode orthogonal to task accuracy: models can produce correct outputs while poisoning the shared state that downstream agents depend on. This isn't about wrong answers; it's about invisible information corruption that compounds across sequential decisions.
This connects directly to the state management gap documented in 'Lost with a Map' (late September), which found that LLMs scatter slot values across context rather than consolidating them. That work explained the architectural reason models struggle with state; this new framework measures the operational consequence when multiple agents rely on that corrupted state. The finding also echoes the co-cheating pattern from 'False Frontiers' (same period), where internal signals diverge from external validation. Here, models pass the immediate task check but fail evidence verification, suggesting similar reward hacking dynamics in collaborative settings.
If the same three models (GPT, Gemini, Qwen) show state corruption rates above 50% when tested on the HELM or SuperGLUE multi-turn subsets within six months, that validates whether this is a general architectural problem or an artifact of the healthcare/disaster scenarios. If not, the framework remains domain-specific.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsGPT · Gemini · Qwen · OffQuery
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as “Right Answers, Wrong States: Hidden Information Failures in Multi-Agent Collaboration”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.