Modelwire
Subscribe

Held-out cross-entropy cannot reliably estimate language model risk

A new theoretical result challenges a foundational assumption in language model evaluation: that held-out cross-entropy loss can reliably estimate true model risk. Researchers prove the estimand is fundamentally inconsistent across possible data-generating distributions and model architectures, meaning no finite sample size guarantees convergence to ground truth. This matters because scaling laws, which guide compute allocation and model selection across the industry, depend on this metric. The inconsistency stems from a topological property where finite and infinite risk states cluster arbitrarily close in model weight space, invisible to any dataset. The finding suggests current benchmarking practices may conflate noise with signal when comparing language models.

Modelwire context

Explainer

The paper doesn't just show cross-entropy is noisy; it proves the metric is fundamentally non-convergent across different data distributions and architectures. This means no amount of held-out data can fix the problem, which is a harder claim than 'benchmarks are unreliable.'

This sits squarely in the evaluation reckoning that's accelerated across recent work. The multidimensional metaphor evaluation framework (August) showed that single-axis metrics miss crucial structure in output quality. The hallucination span detection paper (same week) exposed how binary factuality checks obscure where models actually fail. This cross-entropy result goes deeper: it suggests the statistical foundation of scaling laws themselves may be built on sand. If scaling laws depend on a metric that provably doesn't converge, then model selection and compute allocation decisions across the industry rest on fundamentally unreliable signals.

If major labs release updated scaling law curves using alternative risk estimators (e.g., importance-weighted or stratified holdout methods) within the next six months, that signals they're taking the inconsistency seriously. If no such corrections appear and scaling laws remain unchanged, the field is either dismissing the theoretical result as impractical or hasn't yet grasped its implications for production model selection.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsLanguage models · Cross-entropy loss · Scaling laws

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.LG originally reported this story as Cross-Entropy Risk Estimation for Language Models: Inconsistency Must Be Dense, and the Holdout Method Is No Exception”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.