Autonomous driving models rely on memory more than perception, research finds
Researchers are exposing a blind spot in autonomous driving benchmarks by testing whether models genuinely perceive dynamic scenes or simply memorize static location patterns. By replacing live camera feeds with retrieved memories from prior visits to the same location, they isolate how much driving performance depends on real-time scene understanding versus learned spatial regularities. This work challenges the validity of current evaluation metrics like NAVSIM and Bench2Drive, suggesting that high benchmark scores may overstate a model's ability to handle novel or unexpected traffic situations. The finding has direct implications for safety validation in self-driving systems and raises questions about whether existing benchmarks adequately stress-test decision-making under truly novel conditions.
Modelwire context
Skeptical readThe paper doesn't just show that benchmarks can be gamed; it assumes that swapping live feeds for retrieved memories cleanly separates 'memory' from 'perception.' But this binary framing may be misleading. Humans also rely on spatial priors and learned patterns at familiar locations. The critical omission: whether the models' performance drop under memory injection correlates with actual safety failures in truly novel scenarios, or whether location-based regularities are a legitimate and useful form of scene understanding.
This connects directly to the August 31st work on 'Stress-Testing Efficient Responsible-AI Evaluation,' which found that cost optimization in benchmarking can mask behavioral shifts that matter for safety. Here, the implicit cost is evaluation rigor: NAVSIM and Bench2Drive may be computationally convenient but behaviorally incomplete. Both papers expose the same fault line: aggregate benchmark scores can hide what models actually know. The difference is that the efficiency paper caught this through architectural compression, while this one catches it through data substitution.
If the researchers test their memory-injection method on a held-out set of locations the models never saw during training, and performance remains high, that falsifies their core claim. Conversely, if performance collapses on truly novel locations but holds steady on memorized ones, watch whether NAVSIM and Bench2Drive maintainers adopt location-diversity constraints in their next release. That adoption (or refusal) will signal whether the benchmark community treats this as a real problem or a methodological edge case.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsNAVSIM · Bench2Drive
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. arXiv cs.LG originally reported this story as “Driving on Memory”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.