Vision-language reward models fail paraphrase consistency tests in robotics
Vision-language models deployed as reward functions in robotic learning systems fail to maintain consistency across semantically equivalent instructions, a critical vulnerability for autonomous systems. Researchers introduce ROBORMBENCH, a benchmark of 2,390 real robot trajectories with 21,673 verified paraphrases, demonstrating that both proprietary and open-source VLMs exhibit severe paraphrase fragility. Identical robot behaviors flip between success and failure depending on instruction wording, exposing a fundamental gap between VLM robustness claims and real-world deployment requirements. This finding challenges the assumption that vision-language models are reliable reward arbiters for embodied AI.
Modelwire context
ExplainerThe critical detail the summary underplays: this isn't just a VLM accuracy problem. It's a reward function problem. When the same robot trajectory gets labeled success or failure depending on instruction wording, the learning signal itself becomes corrupted, potentially training robots to overfit to surface-level language patterns rather than actual task semantics.
This connects directly to 'The Visual Insensitivity Gap' from early September, which found VLMs often ignore visual input entirely. ROBORMBENCH reveals a complementary failure: VLMs do process the visual trajectory, but their semantic understanding of language is so brittle that identical behaviors produce contradictory reward signals. Both papers expose the same underlying problem (VLMs conflate surface form with meaning) but from different angles. The paraphrase fragility here also echoes 'Retrieved but not ranked' from the same period, which showed embedders collapse when surface form and semantic structure diverge. For robotics, the stakes are higher: a misaligned reward function doesn't just give wrong answers, it trains the policy in the wrong direction.
If Embodied AI labs (DeepMind Robotics, Boston Dynamics, or academic groups using VLM rewards) publish post-hoc analysis of their trained policies showing overfitting to instruction wording rather than task structure, that confirms ROBORMBENCH's concern is real in production. Alternatively, if a major robotics paper from late 2026 onward switches from VLM rewards back to hand-crafted or learned reward models, that's a market signal the field is moving away from this approach.
Coverage we drew on
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsROBORMBENCH · Vision-language models · VLM reward models
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as “Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.