DUET-DINO: Simultaneous Cross-View World Modeling for Latent Planning in Robot Manipulation
First seen · 9/10/2026, 01:41 AMLatest activity · 9/10/2026, 01:41 AM
Latent world models for robotic manipulation often falter during fine-grained rotational adjustments due to single-camera perspective limits. DUET-DINO couples static side-view cameras with wrist-mounted sensors into a simultaneous cross-view latent model, predicting visual transitions across a full 7-DoF action space. Trained on DROID and RoboArena, the framework achieves a 72.5% success rate on angled reaches and 60% on grasp-and-lift tasks, outperforming dual-view baselines that process feeds independently. The benchmarks also highlight that DINOv3 representations capture local action dynamics significantly better than models like V-JEPA 2.
Event heat · last 24 hours
There are 8 persisted snapshots in the last 24 hours. Peak heat was 0 at 9/12, 14:00; latest heat is 0.