WaiT introduces a wavelet-aware image Transformer for flow matching. It decomposes images into coarse and fine frequency bands, keeping high-frequency bands as noise until coarse structure has emerged, before jointly refining all bands. The paper reports a pixel-space FID of 1.43 on ImageNet at 512×512, up to 50% lower sampling compute, and a 1.3 FID result from its largest 2B model. It also claims strong texture fidelity against latent-space models, scaling to OpenImages and video generation, including an FVD of 0.84 on Kinetics-600 without algorithmic changes.
No heat snapshots are available in the last 24 hours.