Modelwire
Subscribe

OptimismBench reveals systematic bias across 14 of 16 tested language models

Researchers have developed OptimismBench, a novel evaluation framework that exposes systematic directional bias in LLM probability judgments by comparing inverted scenario pairs. The method sidesteps the ground-truth problem that plagues traditional calibration metrics: when a model assigns 70% to success and 15% to failure, the missing 15 points reveal distortion invisible to aggregate scores. Testing across 16 models from 8 providers reveals a striking pattern: 14 exhibit optimistic bias, with pessimism appearing only in Anthropic's frontier models. This finding matters for practitioners deploying LLMs as decision aids, where hidden directional tilt can systematically skew downstream choices in high-stakes domains.

Modelwire context

Explainer

The key insight isn't just that models show directional bias, but that traditional aggregate calibration metrics actively hide it. By comparing paired scenarios (success vs. failure), OptimismBench exposes the 15% of probability mass that standard evaluations miss entirely.

This connects directly to the regional bias framework from late July, which traced how model distortions translate into concrete allocation decisions in hiring and education. OptimismBench operates at a different layer (probability judgment rather than stereotype), but shares the same concern: bias invisible in aggregate metrics becomes visible and harmful once models enter decision-making workflows. The Anthropic finding (pessimism in frontier models only) also echoes the APEX-Accounting benchmark released the same day, which found that even top-tier models fail catastrophically on high-stakes financial tasks requiring perfect accuracy. Together, these suggest that capability and bias are decoupled across the model landscape.

If OptimismBench scores correlate with downstream decision quality in a real-world deployment (e.g., loan approval, clinical triage), that confirms the benchmark predicts harm. If the Anthropic pessimism advantage disappears after instruction tuning or RLHF, that signals the bias is a training artifact rather than a fundamental property of frontier architectures.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsOptimismBench · Anthropic · OpenAI · Google · Meta · Anthropic Claude

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

OptimismBench reveals systematic bias across 14 of 16 tested language models · Modelwire