Hallo4D presents a model-agnostic, training-free framework for reducing spatial and temporal hallucinations in 3D and 4D generation. Its generation-detection-correction pipeline uses large multimodal models to inspect multi-view and multi-frame renderings, summarize duplicated structures, geometric misalignment, jitter, identity flicker, and structural drift, then guide consensus-driven image-space optimization. The framework adds multi-model voting for candidate correction selection, motion-aware keyframe sampling, LMM-guided initialization, appearance alignment, exposure-aware optimization, and visibility pruning. The paper reports improvements over strong baselines across diverse 3D and 4D settings, although the supplied abstract does not provide numerical results.
No heat snapshots are available in the last 24 hours.