Benchmark audit reveals hidden instability in document-grounded LLM tasks
Researchers built Probity, a benchmark of 470 real-world venture-financing documents, to measure how often language models give inconsistent answers to the same question. The critical finding: their own audit uncovered a systematic flaw in excerpt-based benchmarks. When evidence required to answer correctly falls outside the text window shown to the model, instability spikes from 8.7% to 25.5%. This distinction matters because it separates genuine model brittleness from failures caused by incomplete context. The discovery suggests published cross-model agreement metrics may overstate reliability by roughly 20%, forcing a reckoning with how benchmarks are designed and interpreted across the field.
Modelwire context
ExplainerThe real finding isn't that models are unstable on venture documents. It's that the instability metric itself is contaminated by a design choice: when the benchmark excerpt doesn't contain the evidence needed to answer correctly, you're not measuring model reasoning at all, you're measuring what happens when you starve the model of input. This is a meta-benchmark problem, not a model problem.
This connects directly to the multi-model LLM scoring work from last week, which established reliability and validity metrics for educational AI deployment. That paper benchmarked reproducibility across repeated runs; Probity is asking a prior question: are we even measuring what we think we're measuring? Both papers signal growing rigor around what makes a benchmark trustworthy. The verifier error correlation study from Qwen2.5 also fits this pattern, showing that naive statistical assumptions about independence can inflate confidence in evaluation signals. Probity extends that skepticism to the benchmark design layer itself.
If other excerpt-based benchmarks (MMLU, HellaSwag, or domain-specific suites) conduct similar provenance audits and find comparable 15-20 percentage point gaps between in-window and out-of-window performance, that confirms this is a structural issue requiring benchmark redesign across the field. If they don't, Probity's finding may be specific to financial document reasoning.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsProbity
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as “What the Window Does Not Contain: Auditing Provenance in a Document-Grounded Instability Benchmark”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.