Modelwire
Subscribe

Likelihood ranking plateaus while prompting scales across model sizes

Researchers studying 95 decoder-only models from 0.1B to 10B parameters across 10 multiple-choice QA datasets have uncovered a fundamental divergence in how LLMs scale. Likelihood-based ranking of declarative statements shows stable performance across model sizes, while prompted answering improves dramatically with scale. This finding challenges assumptions about evaluation methodology and suggests that how we measure model capability may mask or amplify apparent progress depending on the protocol chosen. The result has implications for benchmarking practices and understanding what scaling actually optimizes.

Modelwire context

Skeptical read

The paper doesn't establish that one method is correct and the other wrong. It documents that two evaluation protocols produce opposite scaling curves on the same models, leaving open whether this reflects genuine capability differences or whether likelihood ranking simply floors out due to task saturation while prompting leaves room for improvement.

This joins a pattern from recent weeks where measurement methodology emerges as the hidden variable. The emoji benchmark audit from late September showed that annotator identity, not model quality, drove 78% of variance in rankings. The code repair study from the same period revealed that 'fixing bugs' and 'rewriting solutions' look identical under standard metrics but represent different capabilities. Here, the same divergence appears: two protocols measuring the same construct produce incompatible conclusions. The common thread is that benchmark design choices can amplify or erase apparent progress independent of what models actually learned.

If the researchers rerun this experiment using calibrated likelihood scores (as in the CORDIAL paper from September) rather than raw model probabilities, and the scaling curves converge, that confirms the finding is a calibration artifact. If they remain divergent after calibration, the result is more likely fundamental to how these architectures optimize different objectives.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsarXiv

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as “Likelihood Ranking doesn't Scale Like Prompting in LLMs”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

Likelihood ranking plateaus while prompting scales across model sizes · Modelwire