Causal Tongue-Tie: LLMs Can Encode Causal Direction, But Their Yes/No Outputs Fail to Express
Source published ·Modelwire updated
Original coverage: arXiv cs.CL ↗·How Modelwire adds context

The development
Researchers have identified a critical gap between what LLMs internally represent about causal relationships and what they output verbally. Using linear probes on hidden states, they recovered near-perfect causal reasoning (97% accuracy) on anti-commonsense questions, yet the models' Yes/No responses collapsed to random performance. This 'Causal Tongue-Tie' reveals that benchmark failures may mask genuine internal understanding, while successes may reflect surface pattern-matching rather than causal cognition. The finding undermines confidence in output-only evaluations and suggests that assessing LLM reasoning requires probing beyond final tokens to distinguish between encoding deficits and expression failures.
Modelwire’s AI-generated summary of coverage from arXiv cs.CL.
Modelwire analysis
ExplainerOur AI-generated reading of the wider context and the next developments to watch.
The deeper provocation here is directional: if linear probes on hidden states can recover causal structure that verbal outputs cannot, then the standard practice of treating benchmark scores as proxies for internal reasoning is not just imprecise but potentially inverted. A model could score well on causal benchmarks through surface pattern-matching while a model that scores poorly might actually encode the correct causal structure.
This connects directly to the surface-versus-semantic noise study covered the same day ('When Do LLM Agents Treat Surface Noise Differently'), which found that meaning-altering perturbations shift model outputs nearly 20 percentage points more than cosmetic changes. Both papers are converging on the same uncomfortable finding from different angles: what a model outputs is a poor and sometimes misleading signal of what it has internally computed. Together they build a case that output-only evaluation is structurally inadequate, not merely noisy.
The critical next step is whether probing methods like linear decoding on hidden states get incorporated into a major public benchmark suite within the next 12 months. If CLadder or a successor adopts probe-based scoring alongside Yes/No accuracy, that would confirm the field is treating this as a measurement problem rather than a curiosity.
This interpretation is generated from the summary above and available source metadata. Our methodology · Report an error
MentionsLLMs · CLadder · linear probe
How this coverage is produced
Modelwire uses AI to generate summaries and context from source headlines, snippets, and selected archive coverage. Automated checks do not verify every claim, and items are not routinely reviewed by a person before publication. Zacaria Solis operates the site. Read the linked source for the full evidence and report errors through our corrections process.
Modelwire summarizes, we don’t republish. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.