Vision-language models for chest radiography do not always need the image

A new causal audit methodology reveals that leading vision-language models trained on chest radiographs may rely heavily on text priors rather than visual analysis. Researchers tested nine systems by occluding image regions and swapping scans, finding that text-only baselines matched multimodal performance within 5.7 percentage points, and a 119B parameter model showed no statistical advantage over a 7B text-only variant. This challenges the field's assumption that strong benchmark scores on medical imaging tasks reflect genuine visual reasoning, raising questions about model transparency in high-stakes clinical deployment and the validity of current evaluation standards.
Modelwire context
ExplainerThe deeper issue isn't that these models are bad at vision; it's that current benchmarks weren't designed to distinguish genuine visual reasoning from sophisticated text-pattern matching. A model can score well by learning that certain clinical phrases predict certain diagnoses, regardless of what the scan actually shows.
This connects to a recurring theme in recent coverage: the gap between what models appear to do and what they actually do internally. The 'Blind Recovery of Latent Domains via Unsupervised Symmetry Discovery' paper from the same day addresses a related problem in a different register, asking how we recover the true underlying signal when the measurement process itself is corrupted or opaque. Here, the corruption is conceptual rather than mathematical: the evaluation signal looks clean but is measuring the wrong thing. Neither story is directly about medical AI, but together they point toward a broader audit problem across applied ML, where surface performance masks structural ambiguity about what a model has actually learned.
Watch whether any of the nine audited systems publish updated evaluation protocols that include text-only ablations as a required baseline. If major clinical AI vendors adopt that standard within the next 12 months, the field is self-correcting; if they don't, regulatory bodies will likely impose it.
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsVision-language models · Chest radiography · Multimodal models · Text-only baseline
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.