
Hallucination in World Models is Predictable and Preventable
Researchers have mapped the failure modes of visual world models, showing that hallucinations cluster predictably in underrepresented regions of the state-action space rather than occurring randomly. The team introduces MMBench2, a 427-hour benchmark with ground-truth dynamics and live simulators, and identifies three distinct hallucination types (perceptual, action-marginalized, scene-diverging) tied to specific pipeline stages. This work shifts world model reliability from an unsolved mystery to an engineerable problem, enabling practitioners to detect and mitigate failures before deployment. The findings matter for embodied AI, robotics, and any system relying on learned environment simulators for planning.62




























