Modelwire
Subscribe

Vision-language models confabulate medical diagnoses based on patient demographics

Illustration accompanying: Hearsay: Vision-Language Medical Diagnoses Without an Image

Frontier vision-language models systematically confabulate medical diagnoses when given only patient demographics and no image, revealing structural bias baked into their reasoning. Claude Opus-4.7, GPT-5.4, and Gemini-3.1-Pro each show predictable diagnostic drift tied to age, race, and gender, with outputs explicitly citing demographic factors as diagnostic evidence. This finding exposes a critical failure mode in high-stakes domains: these models don't abstain under uncertainty but instead generate plausible-sounding false positives anchored to protected characteristics. The result challenges deployment assumptions in healthcare and signals that capability scaling alone does not solve alignment or fairness in specialized reasoning tasks.

Modelwire context

Explainer

The paper's sharpest finding isn't just that bias exists, it's that these models actively fill an evidential vacuum with demographic proxies rather than flagging the absence of input as a reason to abstain. That's a distinct failure from ordinary hallucination: the model is doing something that looks like reasoning.

This connects directly to two threads in recent coverage. The 'Evaluating Regional Bias in LLMs From Abstract Stereotype to Concrete Social Decision-Making' paper established that stereotype leakage reshapes real allocation decisions, not just model outputs, and the medical context here is a sharper version of that same harm pathway. Separately, 'Same Evidence, Different Target' showed that models can appear to reason causally while actually pattern-matching on surface features, which is precisely the mechanism at work when a model cites age or race as diagnostic evidence in the absence of imaging data. Together, these three papers sketch a consistent picture: frontier models substitute familiar statistical patterns for genuine inferential discipline when inputs are incomplete.

Watch whether any of the three named labs (Anthropic, OpenAI, Google) issue updated system-card guidance or deployment restrictions for clinical vision-language applications within the next 90 days. Silence from all three would suggest the finding is being absorbed slowly, which itself is worth flagging.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsClaude Opus-4.7 · GPT-5.4 · Gemini-3.1-Pro · Anthropic · OpenAI · Google

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as Hearsay: Vision-Language Medical Diagnoses Without an Image”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

Vision-language models confabulate medical diagnoses based on patient demographics · Modelwire