Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

Read original
arXiv·Nisarga Nilavadi·Sep 9, 2026, 5:41 PM

DUET-DINO: Simultaneous Cross-View World Modeling for Latent Planning in Robot Manipulation

Papers78

Latent world models for robotic manipulation often falter during fine-grained rotational adjustments due to single-camera perspective limits. DUET-DINO couples static side-view cameras with wrist-mounted sensors into a simultaneous cross-view latent model, predicting visual transitions across a full 7-DoF action space. Trained on DROID and RoboArena, the framework achieves a 72.5% success rate on angled reaches and 60% on grasp-and-lift tasks, outperforming dual-view baselines that process feeds independently. The benchmarks also highlight that DINOv3 representations capture local action dynamics significantly better than models like V-JEPA 2.

Why it's worth reading

It resolves a critical limitation in robotic latent world models—handling 7-DoF rotation and fine spatial control—by pairing wrist-view dynamics with DINOv3 representations for reliable zero-shot planning.

Tags

机器人世界模型具身智能DUET-DINODINOv3潜在规划计算机视觉

Score breakdown

  • Novelty16
  • Impact15
  • Practicality15
  • Credibility15
  • Timeliness17