GraphVid introduces a graph-conditioned image-to-video model for controlling interactions among multiple objects through structured interaction graphs rather than manually drawn trajectories. The authors also present GraphVid-Bench, an interaction-centric video dataset with relational annotations. According to the abstract, GraphVid uses substantially less training data and fewer trainable parameters than prior motion-control approaches, while improving both controllability and video quality against Motion-I2V: FID decreases by up to 39.9%, FVD by 37.6%, PSNR rises from 9.87 to 15.98, and SSIM from 0.38 to 0.61.
No heat snapshots are available in the last 24 hours.