New benchmark reveals multimodal LLMs struggle with hidden social dynamics in meetings

Researchers have released MeetingToM, a benchmark that stress-tests multimodal LLMs on social reasoning in group settings, specifically targeting phenomena like pseudo-consensus where participants mask disagreement under social pressure. This work exposes a critical gap in current model evaluation: existing benchmarks measure overt, easily verifiable signals, but miss the latent group dynamics and unspoken tensions that define real meetings. For practitioners deploying LLMs in workplace collaboration tools, the benchmark signals that current systems lack the nuanced social understanding needed to accurately model participant intent and belief states across distributed audio and visual cues.
Modelwire context
ExplainerThe benchmark's specific focus on pseudo-consensus, the gap between what participants say and what they believe under social pressure, is what separates MeetingToM from prior ToM work, which typically tests dyadic or scripted interactions rather than the emergent, politically charged dynamics of real group settings.
This connects directly to the evaluation gap theme running through recent coverage. The GAMUT benchmark piece from the same day identified a parallel blind spot: current evals catch false claims but miss omissions. MeetingToM extends that critique into the social reasoning domain, where the failure mode is not factual incompleteness but misread intent. Both papers argue that production-grade evaluation must model latent structure, not just surface outputs. The 'Agents in the Wild' piece is also relevant here: if agentic systems are entering workplace collaboration contexts, as that tutorial describes, then the absence of reliable multi-party social reasoning is not an academic gap but an operational liability.
Watch whether any of the major multimodal model labs (Google, OpenAI, Meta) cite MeetingToM in upcoming model cards or evaluation suites within the next two quarters. Adoption there would signal the benchmark has traction beyond the academic circuit.
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsMeetingToM · Multimodal LLMs · Theory of Mind
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as “MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.