PixWorld proposes a unified pixel-space diffusion model for both 3D scene reconstruction and generation. Instead of defining diffusion over latent features, it supervises the process directly through rendered images, avoiding information loss from latent encoders and eliminating the need for a pretrained VAE or RAE. The paper also introduces a geometry perception loss based on features from a pretrained 3D foundation model, adding structural supervision beyond photometric and perceptual image losses. According to the abstract, PixWorld consistently outperforms prior latent-space generation methods and reaches state-of-the-art reconstruction performance.
No heat snapshots are available in the last 24 hours.