Modelwire
Subscribe

Who Checks the Citations? Benchmarking Legal Hallucination Detection

Illustration accompanying: Who Checks the Citations? Benchmarking Legal Hallucination Detection

Legal AI hallucination remains a persistent, worsening problem despite model improvements and professional consequences. Researchers benchmarked five models' ability to automatically detect fabricated citations in court filings, finding that GPT-5 achieves 82.8% recall but only 60.5% F1 score. The work introduces a taxonomy and dataset grounded in over 1,000 real cases of citation fabrication filed year-over-year, establishing that neither scaling nor sanctions have solved the core reliability gap. This signals that downstream verification tools, not just better base models, are becoming essential infrastructure for high-stakes AI deployment.

Modelwire context

Analyst take

The recall-precision gap buried in the numbers is the actual finding: GPT-5 catches most fabricated citations but generates enough false positives that a human reviewer still can't trust the output and walk away. That asymmetry matters enormously for any vendor selling legal AI as a time-saving tool, because noisy alerts erode the productivity case faster than missed ones.

This connects directly to the coherence illusions paper covered the same day ('When Context Misleads'), which showed that surface fluency masks semantic failure in ways models themselves cannot reliably detect. Both papers are pointing at the same structural problem from different angles: models that produce confident, well-formed output are not the same as models that produce accurate output. Together they reinforce that verification infrastructure is not a nice-to-have layer on top of capable base models, it is load-bearing. The legal domain just makes the consequences legible because courts leave paper trails.

Watch whether any of the major legal AI vendors (Thomson Reuters, LexisNexis, Harvey) ship a dedicated citation-verification module citing this taxonomy within the next two quarters. If they do, this benchmark becomes the de facto evaluation standard; if they build proprietary alternatives instead, it signals the research community and the product layer are still operating on separate tracks.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsGPT-5 · arXiv

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

Who Checks the Citations? Benchmarking Legal Hallucination Detection · Modelwire