Claude Haiku solves Ukrainian bar exam without reading questions
Researchers discovered that Claude Haiku 4.5 can solve roughly 38% of Ukrainian judicial exam questions by selecting answers without reading the prompt, revealing systematic option bias in multiple-choice benchmarks. The team identified 11.8% of items where the model answers identically across all option orderings, far exceeding random chance. After filtering out items with detectable legislative text leakage, the benchmark shrinks from 11,990 to 8,128 valid items. This work exposes a critical flaw in how legal and professional benchmarks validate model reasoning: they measure option recognition rather than genuine comprehension, forcing the field to reconsider what multiple-choice scores actually certify about model capability.
Modelwire context
ExplainerThe paper isolates a specific failure mode: models can exploit statistical regularities in how answer options are arranged, independent of question content. This is distinct from data leakage (which the authors also found) and reveals that even 'clean' benchmarks may measure test-taking heuristics rather than legal reasoning.
This connects directly to the latent structure analysis from mid-August, which questioned whether human-normed exams measure the same constructs in LLMs as in humans. The Ukrainian judicial exam work provides a concrete mechanism: models achieve high scores through option recognition patterns that wouldn't transfer to open-ended tasks. The finding also extends the multilingual benchmark concerns raised by L3Cube-IndicQuest and BengaliMCQ, suggesting that multiple-choice formats themselves may be systematically biased toward surface-level matching across languages and domains, not just geographies.
If Anthropic or other labs retest Claude Haiku 4.5 on the 8,128 'clean' items (post-leakage filtering) and performance drops below 25%, that confirms option bias is the primary driver of the 38% baseline. If performance holds above 32%, the bias explanation weakens and suggests genuine legal knowledge is present.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsClaude Haiku 4.5 · Anthropic · UA-JudgeExam · Ukraine Higher Qualification Commission of Judges
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as “Gated Against One Model, Open to the Next: Option-Only Solvability in Legal Multiple-Choice Benchmarks”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.