LLM physician recommendations show reputation dominates, demographics matter
Researchers conducted a large-scale causal audit of seven LLMs, including GPT-4o-mini, to measure what drives physician recommendations when patients use AI assistants for healthcare decisions. Testing 40,000+ choice scenarios with randomized physician profiles, they isolated how reputation signals, demographics, and name-based ethnicity cues influence model outputs. The work exposes a critical gap in AI governance: LLMs now silently shape access to professional services at scale, yet their selection criteria remain opaque and potentially biased. This has immediate implications for healthcare equity and raises broader questions about algorithmic intermediation in high-stakes domains.
Modelwire context
ExplainerThe paper's core contribution isn't that LLMs are biased (expected), but that they're now acting as invisible gatekeepers to professional services at scale, with no transparency into their selection logic. The causal audit methodology isolates what actually drives recommendations, moving beyond correlation to mechanism.
This connects directly to the August work on frozen model steering and abstention (You Only Pass Once). Both papers identify a reliability problem in deployed LLMs where the model's actual decision-making process diverges from what users assume is happening. Here, physicians and patients assume recommendations are based on qualifications; the audit reveals demographic and reputation signals are doing the work. The broader pattern across recent coverage (Style or Signature, Information Satisfaction) shows that evaluation methods and user assumptions consistently lag behind what models actually optimize for, especially when ground truth is opaque or users lack domain expertise to verify outputs.
If OpenAI or other major LLM providers publish their own audits of physician recommendation behavior within six months and report substantially lower bias than this paper found, that suggests they've already patched the models. If they don't respond or claim the audit methodology is flawed, that signals the problem will persist in production systems and likely spread to other high-stakes domains (legal referrals, financial advisors).
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsGPT-4o-mini · OpenAI · arXiv
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as “Whose doctor does the AI recommend? An algorithm audit of reputation and demographic signals in large language model-assisted physician choice”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.