Modelwire
Subscribe

MLLMs ace visual search but fail to replicate human gaze patterns

Researchers benchmarked multimodal LLMs against human eye-tracking data during visual search tasks, revealing a critical divergence: while models match or exceed human performance on target detection and acquisition speed, their internal gaze processes differ fundamentally from human scanpaths. This finding challenges the validity of attention-alignment metrics commonly used to evaluate whether MLLMs genuinely replicate human visual reasoning or merely achieve similar outputs through different mechanisms. The work has direct implications for using these models as proxies for human cognition and for interpreting saliency-based interpretability claims in vision-language systems.

Modelwire context

Skeptical read

The paper doesn't just show models and humans differ in how they look; it reveals that standard attention-alignment metrics used to validate model interpretability may be measuring the wrong thing entirely. Models can solve the task while their internal search process bears no resemblance to human cognition, which means saliency explanations derived from these systems could be systematically misleading.

This echoes a pattern across recent work: benchmark success is decoupling from actual capability. The clinical error detection paper from August 17th found that F1 scores masked poor pairwise discrimination, and the AudioChaps framework exposed how production systems must bridge standardized metrics and real human judgment. Here, the divergence is even more fundamental: the models aren't just failing at nuance, they're succeeding at the task while operating on alien principles. The implication is that we've been using the wrong evaluation lens entirely.

If researchers retrain MLLMs with auxiliary losses that penalize divergence from human scanpaths (not just task accuracy), watch whether this degrades visual search performance or leaves it unchanged. If performance holds steady despite forcing gaze alignment, that confirms the current models are solving the task through redundant mechanisms. If it tanks, that reveals the divergent gaze is actually load-bearing to their efficiency.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsCOCO-Search18 · MLLMs · multimodal large language models

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as Matched Outcomes, Divergent Gaze: How Foveated MLLMs Search Compared to Humans”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

MLLMs ace visual search but fail to replicate human gaze patterns · Modelwire