The paper presents a cheaper way to improve a pretrained pixel-space diffusion model without retraining its backbone. A lightweight prediction head is attached to an intermediate layer of a frozen diffusion Transformer. The difference between intermediate and final predictions becomes a self-guidance direction during sampling: intermediate layers capture coarse, low-frequency structure, while later layers refine high-frequency details. The head can be trained entirely on model-generated samples, and the abstract reports that generated samples outperform real images for this purpose, particularly for enhancing high-frequency detail. Full benchmark results are not included in the supplied summary.
No heat snapshots are available in the last 24 hours.