
How reliable are LLMs when it comes to playing dice?
A systematic evaluation of eight leading LLMs reveals a sharp competency cliff in probabilistic reasoning. While models excel at canonical probability problems (96% accuracy), performance collapses to 59% on counterintuitive variants designed to expose heuristic shortcuts. The study documents two critical failure modes: token bias, where surface-level reformulations tank scores by over 20%, and prompt injection vulnerability, where embedded suggestions degrade accuracy by up to 34%. No model class showed immunity. This work exposes a fundamental gap between statistical fluency and robust reasoning, with implications for deployment in domains requiring reliable uncertainty quantification.62



























