Modelwire
Subscribe

Vision-language models fail diagnostic consistency tests on brain MRI data

Vision-language models exhibit alarming fragility in clinical settings, with diagnostic predictions flipping in nearly half of cases when medical imaging sequences are reordered or label positions shuffled. This arXiv study exposes a critical gap between benchmark accuracy and real-world reliability: VLMs fail to maintain consistent diagnoses despite identical visual evidence, and textual presentation biases trigger misclassification rates exceeding 67%. For healthcare deployment, the finding signals that current VLM robustness claims mask dangerous vulnerabilities in high-stakes domains where input formatting should be irrelevant to clinical judgment.

Modelwire context

Explainer

The study isolates a specific failure mode: VLMs treat identical clinical evidence as diagnostically different based purely on presentation order or label position, not because of corrupted images or missing data. This is distinct from general accuracy degradation under noise.

This connects directly to the August 2nd finding on subtype robustness and calibration. That work showed models maintain high confidence precisely where accuracy fails on unseen variants. This VLM study reveals a parallel problem in the medical domain: models confidently flip diagnoses when formatting changes, signaling no uncertainty. Both expose the same gap: accuracy metrics and benchmark scores hide context-dependent fragility. The earlier work on medical sycophancy (August 2nd) also applies here, since textual perturbations (label shuffling, reordering) function like conversational pressure, triggering model abandonment of correct reasoning under altered input structure rather than genuine evidence change.

If the same VLM models maintain consistent diagnoses when tested on the same cases presented through a standardized clinical interface (e.g., DICOM viewer with fixed label placement), that confirms the fragility is presentation-dependent rather than inherent to the models. If robustness persists only when input formatting is controlled, it signals that deployment risk is mitigatable through interface design rather than requiring model retraining.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsVision-language models · Brain MRI · Histopathology · Diagnostic robustness

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as Evaluating the Diagnostic Robustness of Vision-Language Models Under Visual and Textual Perturbations”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

Vision-language models fail diagnostic consistency tests on brain MRI data · Modelwire