Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

D2PO: Optimizing Diffusion Samplers via Dynamic Preference

First seen · 7/7/2026, 02:05 PMLatest activity · 7/7/2026, 02:05 PM

D2PO, or Dynamic Direct Preference Optimization, reframes low-NFE diffusion sampler optimization as preference alignment rather than conventional student-teacher regression. It models the sampling policy as an energy-based model so preference comparisons become tractable energy differences, and derives an energy function from the pretrained score network to evaluate both structural consistency and fine-grained detail in perturbed spaces. Its dynamic preference mechanism iteratively improves the preferred samples as the sampler learns, aiming to avoid the texture degradation often seen when low-step students imitate high-step teachers. The abstract reports consistent gains over regression-based schedulers, but provides no numerical results here.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/7, 02:05 PMnot independentRepresentative
    D2PO: Optimizing Diffusion Samplers via Dynamic Preference