Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Latent Reward Registers for Diffusion Preference Alignment

First seen · 8/5/2026, 01:00 AMLatest activity · 8/5/2026, 01:00 AM

This paper introduces Latent Reward Registers, learnable position-free tokens prepended to a frozen Diffusion Transformer to estimate terminal preference from intermediate noisy latents. The readout is designed not to modify the generator’s hidden states or velocity field, producing dense differentiable rewards across denoising. The authors use it in Reward-Gradient On-Policy Distillation (RG-OPD) for training and Reward-Guided Sampling (RGS) for parameter-free inference. They report leading pairwise accuracy at noise level u = 0.8, up to 33x fewer GPU hours than online RL baselines, and state-of-the-art results among evaluated training-free methods.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv8/5, 01:00 AMnot independentRepresentative
    Latent Reward Registers for Diffusion Preference Alignment