
ROVE: Unlocking Human Interventions for Humanoid Manipulation via Reinforcement Learning
ROVE addresses a critical bottleneck in humanoid robot training: how to extract value from imperfect human corrections. Vision-language-action models require post-training refinement, but collecting intervention data from humans controlling complex whole-body systems yields noisy, suboptimal trajectories that traditional imitation learning absorbs uncritically. This work combines a hardware-software pipeline for humanoid intervention collection with a reinforcement learning framework that filters and improves flawed human signals rather than copying them directly. The result matters because it unlocks a scalable path to better robot policies without requiring expert-level human operators, reshaping how embodied AI systems move from simulation to real-world deployment.62

























