Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models

Researchers have exposed a critical blind spot in embodied AI: vision-language-action models fine-tuned on robotics tasks often lose commonsense and factual knowledge from their pretrained foundations, yet this degradation has been hard to measure separately from control failures. Act2Answer, a new evaluation protocol, grounds knowledge benchmarks in physical action by converting multiple-choice questions into tabletop manipulation episodes where agents select answers through object placement. This work matters because it clarifies whether VLA failures stem from forgotten knowledge or poor motor generalization, directly informing how practitioners should architect and train embodied systems for real-world deployment.
Modelwire context
ExplainerThe deeper provocation here is not just that VLAs forget things, but that the field has been measuring the wrong thing entirely: task success rates in robotics conflate two distinct failure modes, and practitioners have had no clean way to tell them apart until now.
Act2Answer belongs to a growing cluster of work questioning whether our evaluation tools actually reflect what we care about in deployed AI systems. The MC Dropout paper covered here ('Confidence is Not Reliability') makes a structurally identical argument in medical imaging: a metric practitioners trust, in that case uncertainty scores, turns out not to catch the failures that matter most. Both papers are essentially audits of auditing methods. The pattern is worth naming: as AI systems get embedded in physical or clinical workflows, the gap between benchmark performance and real reliability is becoming the central research problem, not a footnote.
Watch whether robotics benchmarks like Open X-Embodiment or RoboAgent adopt Act2Answer-style knowledge probes alongside task success rates in their next evaluation cycles. If they do, it signals the field has accepted that knowledge retention is a first-class metric, not a curiosity.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsVision-Language-Action models · Act2Answer · VLM · robotics
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.