Finetuning with Scientific Data Increases Hallucinations: A Multi-domain Factuality Evaluation of LLMs

A new benchmark reveals a counterintuitive risk in the push to specialize LLMs for scientific work: domain-specific fine-tuning actually amplifies hallucination rates compared to general-purpose base models. SciFactCheck evaluates 18 models across five scientific domains using a nuanced framework that distinguishes unverifiability, overclaiming, and attribution errors. This finding challenges the assumption that narrow training improves factuality in high-stakes domains like biomedicine and chemistry, forcing practitioners to reconsider whether specialized models require additional safeguards or architectural changes to maintain reliability.
Modelwire context
ExplainerThe more pointed finding isn't just that fine-tuned models hallucinate more, it's that SciFactCheck distinguishes between three failure modes (unverifiability, overclaiming, and attribution errors), which means the problem isn't monolithic and different domains likely fail in different ways. That granularity is what makes this actionable rather than merely alarming.
This lands directly alongside our coverage of MedHal-Loc (also published June 19), which exposed a parallel problem: medical hallucination detectors that claim to localize errors often can't actually do so faithfully. Together, these two papers sketch a troubling picture for clinical AI deployment, where both the models generating outputs and the detectors meant to catch errors are less reliable than their architectures imply. The radiology report generation work we covered the same day, which introduced precision-recall tradeoffs as a deployment control, looks more prescient in this light: if fine-tuning reliably worsens factuality, then inference-time behavioral controls may be a more tractable path than training-time fixes.
Watch whether any of the 18 evaluated models release updated fine-tuning approaches that specifically target the attribution error category SciFactCheck identifies. If overclaiming rates drop without unverifiability rising, that would suggest the failure modes are separable and addressable independently.
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsSciFactCheck · LLMs
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.