Modelwire
Subscribe

Vision-language models fail instruction adherence in annotation tasks

Researchers have exposed a critical vulnerability in vision-language models used as annotation substitutes: VLMs shift their outputs based on image context even when explicit instructions demand text-only judgment. The MIST benchmark demonstrates that both aligned and misleading images alter roughly 20% of VLM decisions across thirteen models tested, suggesting these systems lack the robustness needed for reliable human replacement in annotation workflows. This finding challenges the assumption that VLMs can faithfully execute constrained reasoning tasks and raises questions about their deployment in quality-control and labeling pipelines where instruction adherence is non-negotiable.

Modelwire context

Explainer

The paper's real contribution isn't that VLMs are influenced by images (that's expected) but that they fail silently: models shift outputs without signaling uncertainty or flagging the conflict between visual and textual instructions, making the failure invisible to quality-control systems that rely on them.

This connects directly to the annotation-budget work from late September, which quantified labeling costs across languages and tasks. That research assumed human annotators or reliable automated systems could handle the workload; this paper reveals that VLM-as-annotator introduces a systematic blind spot. The gender-bias heterogeneity paper from the same period is also relevant: just as bias varies unpredictably across models, this work suggests instruction-following robustness may too, complicating vendor selection for annotation pipelines.

If the researchers test whether confidence scores or attention patterns correlate with the 20% failure rate, that would indicate whether practitioners could at least detect when a VLM is unstable. If no such signal exists, annotation workflows using these models need human spot-check rates substantially higher than current industry practice assumes.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsVision-language models · MIST · Misleading-Image Stress Test

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as “It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

Vision-language models fail instruction adherence in annotation tasks · Modelwire