Vision-language models confabulate medical diagnoses based on patient demographics
Source published ·Modelwire updated
Original coverage: arXiv cs.CL ↗·How Modelwire adds context

The development
Frontier vision-language models systematically confabulate medical diagnoses when given only patient demographics and no image, revealing structural bias baked into their reasoning. Claude Opus-4.7, GPT-5.4, and Gemini-3.1-Pro each show predictable diagnostic drift tied to age, race, and gender, with outputs explicitly citing demographic factors as diagnostic evidence. This finding exposes a critical failure mode in high-stakes domains: these models don't abstain under uncertainty but instead generate plausible-sounding false positives anchored to protected characteristics. The result challenges deployment assumptions in healthcare and signals that capability scaling alone does not solve alignment or fairness in specialized reasoning tasks.
Modelwire’s AI-generated summary of coverage from arXiv cs.CL.
Modelwire analysis
ExplainerOur AI-generated reading of the wider context and the next developments to watch.
The paper's sharpest finding isn't just that bias exists, it's that these models actively fill an evidential vacuum with demographic proxies rather than flagging the absence of input as a reason to abstain. That's a distinct failure from ordinary hallucination: the model is doing something that looks like reasoning.
This connects directly to two threads in recent coverage. The 'Evaluating Regional Bias in LLMs From Abstract Stereotype to Concrete Social Decision-Making' paper established that stereotype leakage reshapes real allocation decisions, not just model outputs, and the medical context here is a sharper version of that same harm pathway. Separately, 'Same Evidence, Different Target' showed that models can appear to reason causally while actually pattern-matching on surface features, which is precisely the mechanism at work when a model cites age or race as diagnostic evidence in the absence of imaging data. Together, these three papers sketch a consistent picture: frontier models substitute familiar statistical patterns for genuine inferential discipline when inputs are incomplete.
Watch whether any of the three named labs (Anthropic, OpenAI, Google) issue updated system-card guidance or deployment restrictions for clinical vision-language applications within the next 90 days. Silence from all three would suggest the finding is being absorbed slowly, which itself is worth flagging.
This interpretation is generated from the summary above and available source metadata. Our methodology · Report an error
MentionsClaude Opus-4.7 · GPT-5.4 · Gemini-3.1-Pro · Anthropic · OpenAI · Google
How this coverage is produced
Modelwire uses AI to generate summaries and context from source headlines, snippets, and selected archive coverage. Automated checks do not verify every claim, and items are not routinely reviewed by a person before publication. Zacaria Solis operates the site. Read the linked source for the full evidence and report errors through our corrections process.
Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as “Hearsay: Vision-Language Medical Diagnoses Without an Image”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.