This paper introduces Latent Reward Registers, learnable position-free tokens prepended to a frozen Diffusion Transformer to estimate terminal preference from intermediate noisy latents. The readout is designed not to modify the generator’s hidden states or velocity field, producing dense differentiable rewards across denoising. The authors use it in Reward-Gradient On-Policy Distillation (RG-OPD) for training and Reward-Guided Sampling (RGS) for parameter-free inference. They report leading pairwise accuracy at noise level u = 0.8, up to 33x fewer GPU hours than online RL baselines, and state-of-the-art results among evaluated training-free methods.
No heat snapshots are available in the last 24 hours.