The paper introduces RynnWorld-4D, a generative world model that jointly predicts future RGB frames, depth maps, and optical flow from one RGB-D observation plus a language instruction. Its RGB-DF representation is designed to connect appearance, geometry, and motion more directly to robotic actions. The authors also release Rynn4DDataset 1.0, containing more than 254.4 million frames from egocentric human and robotic manipulation videos with pseudo-labels for depth and optical flow. RynnWorld-4D-Policy uses the model’s internal 4D representations to predict actions in one forward pass, targeting closed-loop dexterous bimanual manipulation.
No heat snapshots are available in the last 24 hours.