PointDiT presents a minimalist pixel-space Diffusion Transformer for monocular geometry estimation. Built on a plain ViT, it directly models raw 3D point-map patches and conditions generation on image tokens from pretrained DINOv3. The diffusion backbone is trained from scratch, avoiding point-map tokenizers, latent-space compression, and hybrid architectural components. According to the abstract, PointDiT outperforms more complex latent-diffusion approaches while producing sharper geometry and greater robustness in ambiguous regions such as transparent objects. The paper argues that architectural and loss-function complexity may not be necessary for strong single-image 3D reconstruction.
No heat snapshots are available in the last 24 hours.