How reliable are LLMs when it comes to playing dice?

A systematic evaluation of eight leading LLMs reveals a sharp competency cliff in probabilistic reasoning. While models excel at canonical probability problems (96% accuracy), performance collapses to 59% on counterintuitive variants designed to expose heuristic shortcuts. The study documents two critical failure modes: token bias, where surface-level reformulations tank scores by over 20%, and prompt injection vulnerability, where embedded suggestions degrade accuracy by up to 34%. No model class showed immunity. This work exposes a fundamental gap between statistical fluency and robust reasoning, with implications for deployment in domains requiring reliable uncertainty quantification.
Modelwire context
ExplainerThe more alarming finding isn't the accuracy drop itself but what causes it: models aren't reasoning about probability, they're pattern-matching to familiar problem shapes. When surface framing changes while the underlying math stays identical, scores collapse, which means these models have learned to recognize probability problems rather than solve them.
This connects directly to a cluster of failure-mode research Modelwire has been tracking. The financial LLM audit from early June showed that Bitcoin asset preferences shifted dramatically based on framing context alone, which is structurally the same vulnerability documented here: surface reformulation overrides underlying logic. The eating disorder safety study from the same period found that specific linguistic patterns trigger unsafe outputs, another instance where prompt surface features defeat intended model behavior. Taken together, these papers suggest token bias and framing sensitivity aren't edge cases but recurring architectural properties that show up across domains, from clinical safety to portfolio allocation to basic probability.
Watch whether any of the eight evaluated models releases updated reasoning benchmarks in the next two quarters that specifically include counterintuitive probability variants. If none do, that's evidence the field is still optimizing for canonical problem performance rather than addressing the heuristic shortcut problem this paper identifies.
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsLLMs · Chain-of-Thought prompting
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.