Modelwire
Subscribe

LLM accuracy collapses as candidate sets grow, study finds

Researchers have uncovered a critical vulnerability in how LLMs handle decision-making under scale. As the number of candidate options increases, model accuracy drops substantially across different architectures and prompting strategies, contradicting assumptions baked into most current benchmarks. The root cause isn't poor long-context retrieval but rather a systematic collapse in the model's ability to maintain confidence separation between correct and plausible wrong answers. This finding reshapes how practitioners should interpret LLM reasoning capabilities and suggests current evaluation suites may overstate real-world performance in high-cardinality selection tasks.

Modelwire context

Explainer

The paper isolates a specific failure mode: not that LLMs can't find the right answer in a large set, but that they lose the ability to rank it confidently above plausible alternatives. This is distinct from long-context degradation and suggests the problem is architectural rather than informational.

This connects directly to 'The Decomposition Tax' study from late September, which found that LLMs hemorrhage accuracy when intermediate pipeline stages lose problem context. Both papers identify a common thread: LLMs degrade not because they lack information, but because they lose the structural coherence needed to maintain decision confidence across stages or scale. Where decomposition taxes information visibility at boundaries, choice overload taxes confidence separation at scale. Together they suggest the real bottleneck in production systems isn't retrieval or reasoning capacity, but the model's ability to preserve signal-to-noise ratios under complexity.

If the same models tested here show recovery when given explicit ranking instructions or confidence calibration prompts, that confirms the issue is remediable through inference design. If accuracy remains flat regardless of prompting strategy, watch whether practitioners begin building external ranking layers (learned rerankers or symbolic filters) as a standard post-hoc step rather than relying on end-to-end model selection.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsLLMs

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as “Overwhelmed by Choice: Studying LLM Decision Making at Scale”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

LLM accuracy collapses as candidate sets grow, study finds · Modelwire