Random rewards reveal hidden capacity in language models
Researchers propose random-reward reinforcement learning as a diagnostic tool to measure what capabilities lie dormant in pretrained language models. The work reframes a long-standing puzzle: why even nonsensical reward signals can boost LLM performance during RL fine-tuning. Rather than attributing gains to training mechanics or data leakage, the authors argue this phenomenon reveals model reachability, the latent capacity a model can unlock under new constraints without changing its base weights. Using OLMo checkpoints as test cases, they demonstrate that two models with identical benchmark scores can have vastly different potential trajectories. This shifts how researchers should interpret probing results and has implications for understanding what untapped capability exists in deployed models.
Modelwire context
ExplainerThe paper's core claim is that identical benchmark scores can mask radically different capability ceilings. What's missing from the summary: this challenges the assumption that benchmarks measure what models can actually do, not just what they currently do.
This connects directly to 'The Missing Primitive' (October 1st), which also argues that surface-level correctness scores obscure genuine capability gaps. Both papers use diagnostic frameworks to expose the gap between what benchmarks show and what models can unlock under different training regimes. The anisotropy work from late September also fits here: if certain channels drive perplexity but not reasoning, then random rewards might be activating dormant reasoning pathways that benchmarks never probe. Together, these three papers suggest deployed models are systematically underestimated by current evaluation methods.
If researchers apply this random-reward probe to closed-weight models (GPT-4, Claude) and report capability gaps that correlate with downstream task performance on held-out reasoning benchmarks, that validates the reachability concept. If the gaps don't correlate, the phenomenon may be an artifact of open-weight model training rather than a general property of LLMs.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsOLMo
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as “Probe with Participation Trophies: Random-Reward RL as a Probe of LLM Capability”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.