DreamTraj predicts relative 6-DoF object trajectories from a single RGB image and a task instruction, without requiring video, depth, or CAD models at inference. Instead of generating a video and recovering motion afterward, it reads motion-related information from intermediate representations of a frozen image-to-video diffusion model at an early denoising step. A lightweight flow-matching Reader decodes query-key attention tracks and pooled hidden states. The paper also introduces MOVE, a dataset of 5,038 object-centric egocentric trajectories paired with fine-grained natural-language instructions rather than coarse verb-noun labels.
No heat snapshots are available in the last 24 hours.