UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models
UniWorld-View presents a unified framework for controllable novel-view synthesis from monocular images or videos under large camera-baseline changes. It combines explicit 3D guidance from an occlusion-aware point-cloud rendering strategy with video diffusion backbones. The authors state that this design addresses visibility ambiguities, improves geometric consistency, and supports precise camera control even with extremely limited input coverage. The framework can also generate multi-view videos for downstream dynamic 3D Gaussian Splatting reconstruction. Reported evaluations use the WorldScore benchmark and zero-shot novel-view-synthesis benchmarks, although the supplied abstract does not provide numerical results.
Why it's worth reading
Large-baseline view synthesis is moving beyond sparse reconstruction toward hybrid geometric-generative systems. This paper is timely because it connects explicit visibility-aware 3D priors, controllable video diffusion, and downstream dynamic 3DGS reconstruction.