Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Optimizing Visual Generative Models via Distribution-wise Rewards

First seen · 7/3/2026, 12:00 PMLatest activity · 7/3/2026, 12:00 PM

This paper argues that sample-wise rewards in reinforcement learning for visual generation can cause reward hacking, mode collapse, and visual artifacts. It proposes distribution-wise rewards that evaluate generated samples as a set, improving alignment with real-world data distributions. A subset-replace strategy reduces the cost of estimating these rewards by modifying only a small portion of a generated reference set. The authors also use reinforcement learning to optimize post-hoc model-merging coefficients, addressing train-inference inconsistency introduced by stochastic differential equations. Reported FID-50K improves from 8.30 to 5.77 for SiT and from 3.74 to 3.52 for EDM2.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/3, 12:00 PMnot independentRepresentative
    Optimizing Visual Generative Models via Distribution-wise Rewards