DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation
DyPES-VLA targets generalist robot manipulation across heterogeneous embodiments. It trains a vision-language model with a future-prediction objective so shared query representations capture object motion, contact, and interaction-induced scene changes. An embodiment-specific Mixture-of-Experts action head then maps these shared dynamics priors directly into each robot’s native action space, avoiding manual action-format alignment. According to the supplied abstract, the policy reaches 98.0% success on LIBERO, 59.25% on RoboCasa-GR1, and 89.02% on RoboTwin 2.0 across simulation and real-world evaluations.
Why it's worth reading
Cross-embodiment data integration is a central obstacle for generalist robot policies, and this work directly addresses both transferable dynamics representations and incompatible native action spaces with results reported on three benchmarks.