Modelwire
Subscribe

Transformers hide logical reasoning in hidden states despite behavioral failure

Researchers probed whether transformer models truly understand logical reasoning or merely pattern-match their way to correct answers. Testing five open-weight models with controlled premise-claim pairs, they found a striking gap: models performed near chance on behavioral tasks, yet their hidden representations encoded logical validity with near-perfect fidelity. This decoding held across unseen templates, domains, and inference types, even when the model's final answer was wrong. The finding reshapes how we interpret model capabilities. A system can harbor sophisticated internal structure without surfacing it behaviorally, complicating both safety audits and capability claims. Insiders now face a harder interpretability problem: decodability alone doesn't guarantee reasoning.

Modelwire context

Explainer

The paper's core contribution isn't just that models have hidden structure, but that this structure can be perfectly readable yet behaviorally inert. Models scored near chance on actual reasoning tasks while their latent representations encoded logical validity with near-perfect fidelity across held-out cases. This decoupling is the finding.

This directly extends 'Lagged Coupling' from yesterday, which found that readability precedes causal influence in Pythia models. That work showed the lag exists; this paper shows it can be extreme. Where Lagged Coupling demonstrated a temporal ordering, this work reveals that decodability can persist indefinitely without ever becoming behaviorally effective. Both papers converge on the same implication for alignment: mechanistic interpretability tools that rely on linear probes may give false confidence that we understand what a model will actually do. The StateSwap work also fits here, showing that hidden states can encode distinct computational pathways that don't reliably surface in outputs.

If researchers can identify which architectural or training conditions collapse this gap (making latent validity representations causally effective), that's the real lever for improving reasoning. Watch whether follow-up work on this paper attempts causal interventions to flip behavioral performance by steering the hidden representations they decoded, or whether the gap persists even under intervention. That result determines whether the representations are truly 'reasoning' or merely statistical artifacts.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

Mentionstransformer models · language models · logical reasoning · hidden states

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as When Decodability Is Not Enough: Logical Validity Representations, Behavioral Dissociation, and Causal Tests in Language Models”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

Transformers hide logical reasoning in hidden states despite behavioral failure · Modelwire