Benchmark exposes LLM theory-of-mind gap beyond game rules
Researchers have built Avalon-ToM-Bench, a structured evaluation framework that isolates theory-of-mind reasoning in large language models by decomposing it into four distinct cognitive tasks using asymmetric-information game mechanics. Testing 28 LLMs revealed a critical gap: models excel at rule comprehension but fail at genuine social reasoning, suggesting current architectures conflate factual knowledge with the ability to model other agents' beliefs and intentions. This finding reshapes how the field should measure and train for multi-agent reasoning, moving beyond end-to-end gameplay success toward diagnostic isolation of reasoning failures.
Modelwire context
ExplainerThe critical insight isn't that models fail at Avalon, but that they fail in a specific way: they memorize rules and social facts without actually modeling other agents' mental states. This distinction between factual knowledge and genuine belief reasoning is what the benchmark isolates.
This work sits directly alongside the social judgment benchmarking from August 10th (the persona ratings study), which found that LLMs show internal consistency on social assessments but may not be substituting for human judgment in meaningful ways. Avalon-ToM-Bench goes deeper: it suggests the consistency itself might be surface-level pattern matching rather than genuine social reasoning. Both papers flag the same underlying problem from different angles: we've been measuring what models can retrieve or rank, not what they can actually infer about other minds. The PragMatch work on vision-language models reinforces this pattern, showing that multimodal systems also mistake shortcut cues for genuine reasoning.
If the same 28 models show better performance when given explicit access to other players' previous statements (removing the need to infer intent), that confirms the gap is reasoning-specific rather than architectural. Watch whether follow-up work on the benchmark shows whether chain-of-thought prompting or in-context examples of belief modeling close the gap, or whether the failure persists even with scaffolding.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsAvalon-ToM-Bench · The Resistance: Avalon · LLMs
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as “Avalon-ToM-Bench: Evaluating Fine-Grained Theory of Mind via Asymmetric Game Mechanics”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.