D2PO, or Dynamic Direct Preference Optimization, reframes low-NFE diffusion sampler optimization as preference alignment rather than conventional student-teacher regression. It models the sampling policy as an energy-based model so preference comparisons become tractable energy differences, and derives an energy function from the pretrained score network to evaluate both structural consistency and fine-grained detail in perturbed spaces. Its dynamic preference mechanism iteratively improves the preferred samples as the sampler learns, aiming to avoid the texture degradation often seen when low-step students imitate high-step teachers. The abstract reports consistent gains over regression-based schedulers, but provides no numerical results here.
No heat snapshots are available in the last 24 hours.