RynnWorld-Teleop proposes “digital teleoperation,” replacing a physical robot with a generative, robot-centric world model. An operator’s hand-pose stream conditions egocentric video generation from a single reference image, while the pose stream remains an embodiment-agnostic action label that can be retargeted to different robots. The system combines depth-aware skeletal conditioning, progressive human-to-robot training on a video Diffusion Transformer, and streaming autoregressive distillation. The abstract reports over 40 FPS on one H100 GPU, zero-shot Sim2Real transfer for dexterous and bimanual tasks, and consistent gains when synthetic trajectories augment real-world datasets.
No heat snapshots are available in the last 24 hours.