This paper presents a two-stage approach to video motion transfer when reference and target objects have substantially different morphology, articulation, or deformation mechanisms. Stage I learns complementary multi-granularity abstract motion views and uses them to bootstrap cross-category video pairs whose dynamics remain transferable. Stage II distills that supervision into direct reference-video-conditioned generation, eliminating explicit motion extraction at inference. The authors also introduce OpenVMT-Dataset and OpenVMT-Bench for image- and text-conditioned evaluation across Same, Near, and Far category gaps. The abstract reports state-of-the-art motion fidelity and target preservation, while dataset release is planned upon acceptance.
No heat snapshots are available in the last 24 hours.