This paper studies register tokens in Diffusion Transformers (DiTs). Unlike Vision Transformers, DiTs do not show high-norm patch-token outliers, yet they still benefit from registers. The effect is stronger in pixel-space DiTs than in latent-space DiTs. Analysis of intermediate representations suggests that registers produce cleaner feature maps at high noise levels, potentially supporting better visual structure and coherence during generation. The authors also observe that recent pixel-space DiT designs implicitly include register-like mechanisms, and propose Register Guidance to amplify the contribution of registers during sampling or generation.
No heat snapshots are available in the last 24 hours.