DUET-DINO: Simultaneous Cross-View World Modeling for Latent Planning in Robot Manipulation
Latent world models for robotic manipulation often falter during fine-grained rotational adjustments due to single-camera perspective limits. DUET-DINO couples static side-view cameras with wrist-mounted sensors into a simultaneous cross-view latent model, predicting visual transitions across a full 7-DoF action space. Trained on DROID and RoboArena, the framework achieves a 72.5% success rate on angled reaches and 60% on grasp-and-lift tasks, outperforming dual-view baselines that process feeds independently. The benchmarks also highlight that DINOv3 representations capture local action dynamics significantly better than models like V-JEPA 2.
Why it's worth reading
It resolves a critical limitation in robotic latent world models—handling 7-DoF rotation and fine spatial control—by pairing wrist-view dynamics with DINOv3 representations for reliable zero-shot planning.