HarmoHOI presents a unified diffusion framework for synchronized multi-view hand-object interaction videos and globally aligned 3D point tracks. It introduces a Mixture of Multi-view Diffusion Transformer that jointly models RGB videos and point tracks, representing the latter as pseudo-videos to reduce the gap between 3D geometry and 2D foundation-model latents. A Global Motion Aligning Diffusion module refines coarse tracks into metric-scale, globally aligned trajectories. The authors also use a hybrid curriculum to transfer priors from single-view data to scarce multi-view HOI data. The abstract reports state-of-the-art visual quality, motion plausibility, and geometric consistency, without listing quantitative results.
No heat snapshots are available in the last 24 hours.