UDT introduces a U-Net-shaped diffusion transformer that performs downsampling and upsampling through data-adaptive token merging while preserving the DiT token dimension. According to the submitted abstract, its XL model surpasses SiT’s 7.9 FID after 1,400 epochs without classifier-free guidance in only 40 epochs on 256×256 ImageNet, described as roughly 40× faster convergence. With CFG, the reported FID reaches 1.38 after 320 epochs using SD-VAE and 1.35 after 500 epochs using VA-VAE. The full experimental setup and reproducibility details still require verification.
No heat snapshots are available in the last 24 hours.