WorldSculpt: Generating Compositional Worlds from Grounded Videos
Reconstructing heavily occluded, cluttered scenes typically yields a monolithic mesh with missing geometry in hidden areas. WorldSculpt circumvents this limitation by equipping a single-object generative prior (Pixal3D) with a multi-view conditioning pathway to ground generation in posed video observations. Trained entirely on isolated objects in canonical space without scene-level supervision, the method decomposes complex environments into hundreds of standalone object meshes and successfully converts raw 3D Gaussian Splatting captures into modular, interactive assets.
Why it's worth reading
It demonstrates a scalable path to decompose complex, cluttered environments into clean, individual object meshes using only single-object generative priors, avoiding the need for scene-level 3D supervision.