Read original
HuggingFace Daily PapersHaiyang ZhouPapers84

UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models

UniWorld-View presents a unified framework for controllable novel-view synthesis from monocular images or videos under large camera-baseline changes. It combines explicit 3D guidance from an occlusion-aware point-cloud rendering strategy with video diffusion backbones. The authors state that this design addresses visibility ambiguities, improves geometric consistency, and supports precise camera control even with extremely limited input coverage. The framework can also generate multi-view videos for downstream dynamic 3D Gaussian Splatting reconstruction. Reported evaluations use the WorldScore benchmark and zero-shot novel-view-synthesis benchmarks, although the supplied abstract does not provide numerical results.

Why it's worth reading

Large-baseline view synthesis is moving beyond sparse reconstruction toward hybrid geometric-generative systems. This paper is timely because it connects explicit visibility-aware 3D priors, controllable video diffusion, and downstream dynamic 3DGS reconstruction.

Tags

novel-view-synthesisvideo-diffusion3D-guidanceocclusion3DGScamera-controlWorldScore