Modelwire
Subscribe

Transformers fail two-hop reasoning when patterns shift, mechanistic study shows

Researchers have identified a critical failure mode in transformer reasoning: models trained on symbolic tasks generalize reliably when new queries follow seen patterns, but collapse entirely when the second reasoning step deviates from training data. Through mechanistic analysis, the team traced this brittleness to the absence of stable entity representations across different contexts. This finding challenges assumptions about multi-hop reasoning in LLMs and suggests that apparent reasoning capability may rest on shallow pattern matching rather than compositional understanding. The work has direct implications for evaluating and improving model robustness in knowledge retrieval and logical inference tasks.

Modelwire context

Explainer

The paper isolates a precise failure mode: models don't fail because they can't do multi-hop reasoning in principle, but because they lack stable internal representations of entities across different contexts. This distinction matters because it suggests the problem isn't reasoning depth but representation consistency.

This connects directly to the spatial reasoning work from earlier today (Geo-Spatial Concept Probing), which also found that LLMs pattern-match rather than genuinely compose concepts. Both papers use controlled benchmarks to expose where models collapse under compositional pressure. The current work goes further by identifying the mechanistic culprit (unstable entity representations) rather than just documenting the failure. This builds on a pattern in recent coverage: when you isolate one cognitive property at a time (compositionality here, abstraction and grounding in the spatial work), models reveal sharp blindspots that aggregate benchmarks typically obscure.

If researchers can stabilize entity representations through targeted training and the same models then generalize to novel second-hop patterns, that confirms the diagnosis. If generalization still fails despite stable representations, the root cause lies elsewhere and the mechanistic explanation here is incomplete.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsTransformers · Language models

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as Why Knowing Both Hops Is Not Enough: Understanding Two-Hop Generalization in Language Models”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

Transformers fail two-hop reasoning when patterns shift, mechanistic study shows · Modelwire