Semantic context, not just surprise, drives brain responses to speech
Researchers demonstrate that semantic relevance, not just word-level surprise, predicts neural activity during language comprehension. By modeling how incoming words relate to recent semantic context using fMRI data from two naturalistic speech datasets, the work challenges the dominance of surprisal-based models in neurolinguistics. This finding matters for language model interpretability: it suggests that capturing contextual coherence alongside local probability distributions may better align AI systems with human language processing, informing both cognitive modeling and next-generation architectures that aim to replicate naturalistic comprehension.
Modelwire context
ExplainerThe paper doesn't just show semantic relevance predicts brain activity; it does so while controlling for surprisal, meaning the two are separable predictors. This matters because surprisal-based models have dominated neurolinguistics for over a decade, and this is direct evidence that they're incomplete.
This connects to the earlier finding on formal semantic structure and human disagreement (the NLI paper from July 17). Both papers treat semantic organization as a measurable, constraining force on cognition rather than noise. Where that work showed semantic monotonicity predicts annotation variance, this one shows semantic coherence predicts neural firing patterns. Together they suggest the field is converging on treating semantic structure as a primary explanatory variable, not a secondary gloss on probability.
If downstream language model work (particularly interpretability papers over the next 6 months) begins explicitly modeling contextual semantic coherence as a separate loss term from next-token prediction, that confirms this finding is reshaping how researchers think about alignment between models and human processing. If papers continue treating surprisal as sufficient, the result remains confined to neuroscience.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsAlice dataset · Moth dataset · fMRI BOLD · GAMMs
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as “Contextual Semantic Relevance Tracks fMRI BOLD Responses During Naturalistic Speech Comprehension”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.