DistillAlign revisits autoregressive video distillation pipelines that separate initialization from Distribution Matching Distillation (DMD), arguing that high visual quality alone can hide poor distribution coverage. The paper introduces a shared-latent-space protocol for measuring precision and coverage between student and teacher distributions. It also proposes joint distillation, combining DMD’s mode-seeking reverse-KL objective with a Consistency Distillation-based mode-covering constraint. According to the reported experiments, the method improves generation quality, coverage, and diversity, and a Wan-1.3B DMD teacher with DistillAlign outperforms baselines refined with Wan-14B.
No heat snapshots are available in the last 24 hours.