Modelwire
Subscribe

Vision-language models ignore visual input on majority of benchmark samples

Researchers have identified a critical blind spot in vision-language models: up to 97% of test samples show no measurable change in model outputs when question-relevant visual regions are obscured, suggesting these systems often ignore their visual input entirely. The Visual Insensitivity Gap appears consistent across different VLM architectures, indicating the problem stems from dataset properties rather than model design. This finding undermines confidence in aggregate benchmark scores and raises questions about whether current evaluation methods actually validate multimodal reasoning or merely language performance.

Modelwire context

Explainer

The paper doesn't just document that VLMs ignore visual input on some tasks; it isolates a measurement problem: existing benchmarks may reward language-only performance while appearing to validate multimodal reasoning, making it impossible to know whether reported improvements reflect actual visual reasoning or just better text generation.

This connects directly to the BenchMIRT investigation from Hugging Face, which exposed how most benchmarks measure narrow task performance rather than genuine reasoning. The Visual Insensitivity Gap is a concrete instantiation of that broader problem: a metric (aggregate benchmark score) that looks valid but systematically misses a critical failure mode. The finding also echoes the post-hoc alignment work on LLM judges, which showed that optimizing against collapsed ground truth obscures the actual signal in evaluation data. Both papers argue that current evaluation infrastructure is measuring the wrong thing.

If researchers rerun standard VLM benchmarks (COCO Captions, VQA v2, TextVQA) using the Visual Sensitivity Index as a filter and report what fraction of improvements disappear when visual insensitivity is controlled for, that will confirm whether benchmark gains over the past 18 months reflect genuine multimodal progress or dataset artifacts. If the number is above 50%, the field has been chasing a mirage.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsVision-language models · Visual Insensitivity Gap · Visual Sensitivity Index

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.LG originally reported this story as The Visual Insensitivity Gap: Diagnosing When Vision-Language Models Fail to Use Visual Evidence”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

Related

Closing CLIP's modality gap can harm zero-shot accuracy

arXiv cs.CL·

Cultural signals push language models into wrong frameworks, study finds

arXiv cs.LG·

Efficient document VLM matches human annotation costs in regulated workflows

arXiv cs.CL·
Vision-language models ignore visual input on majority of benchmark samples · Modelwire