Modelwire
Subscribe

Anisotropy in LLMs serves distinct roles in perplexity versus reasoning

Researchers have identified a counterintuitive role for anisotropy, the concentration of large activations in a small fraction of residual channels that emerges during LLM post-training. Rather than treating it as a defect, this work reveals that roughly 5% of channels are critical for language modeling performance, yet contribute minimally to reasoning quality. The finding exposes a fundamental asymmetry: supervised fine-tuning reshapes these outlier channels while reinforcement learning leaves them largely untouched. This distinction between what drives perplexity and what drives reasoning accuracy has implications for how practitioners design post-training pipelines and optimize model efficiency.

Modelwire context

Explainer

The paper's core contribution is not just identifying anisotropy, but showing it's functionally asymmetric: the same concentrated channels that hurt efficiency barely affect reasoning performance, meaning you cannot simply prune them without collateral damage to language modeling.

This connects directly to the spike brittleness work from late September on vision-language models, which also uncovered how a small number of channels can dominate computation in ways that look like defects but may serve specific functions. Here, anisotropy in text LLMs plays a similar dual role. The key difference is that supervised fine-tuning and reinforcement learning reshape these channels differently, suggesting post-training method choice is not just about final capability but about which internal pathways get reinforced. This distinction echoes the routing drift paper from the same period, which showed that parameter reassignment alone does not diagnose failure in merged models. Like that work, this paper asks practitioners to look deeper than surface-level metrics.

If practitioners report that pruning the identified 5% of anisotropic channels while keeping the remaining 95% intact preserves reasoning benchmarks (MATH, ARC-Challenge) but degrades perplexity by less than 2%, that confirms the asymmetry claim. If instead pruning causes reasoning drops comparable to perplexity drops, the finding's practical value for efficiency work collapses.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsLLM · supervised fine-tuning · reinforcement learning · residual channels

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.LG originally reported this story as “Understanding and Exploiting Anisotropy in Post-Training”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

Anisotropy in LLMs serves distinct roles in perplexity versus reasoning · Modelwire