Modelwire
Subscribe

Researchers measure language model forecast coherence via arbitrage detection

Researchers have developed a formal method to measure whether language model probability estimates form a coherent belief system, using Dutch-book arbitrage as a metric. The approach applies de Finetti's theorem to detect when LLM forecasts contain exploitable inconsistencies, without requiring ground-truth labels. This addresses a critical gap: as users increasingly rely on LLMs for probabilistic reasoning about life decisions and economic outcomes, the internal consistency of model-generated probabilities directly affects decision quality. The work signals growing scrutiny of LLM reliability beyond accuracy metrics, focusing instead on whether models maintain logically sound probability distributions.

Modelwire context

Explainer

The paper's core contribution is detecting probability incoherence without ground truth labels. Most LLM evaluation requires comparing outputs against a known correct answer; this method instead checks whether the model's own probability assignments contradict each other, making it applicable to open-ended forecasting where no ground truth exists.

This work sits squarely in the recent wave of papers questioning what current LLM evaluation actually measures. Like BenchMIRT's critique of narrow task performance and the post-hoc alignment paper's pivot from collapsed ground truth to distribution matching, this research reframes the evaluation target entirely. Rather than asking 'did the model pick the right answer,' it asks 'are the model's beliefs internally consistent.' The timing matters: as the field has spent the last month (early September 2026) exposing gaps in how we measure LLM behavior, this paper offers a practical tool for one specific gap that matters for real-world decision-making.

If practitioners deploying LLMs for financial forecasting or medical triage adopt Dutch-book testing as a pre-deployment filter in the next six months, that signals the field is moving beyond accuracy metrics toward coherence checks. Conversely, if the method remains confined to academic benchmarking without industry adoption by end of Q1 2027, it suggests the practical overhead or false-positive rate makes it unworkable at scale.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

Mentionsde Finetti · language models

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.LG originally reported this story as Dutch Books for Language Models”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

Researchers measure language model forecast coherence via arbitrage detection · Modelwire