The paper presents a feed-forward framework that decomposes unposed multi-view images into instance-structured 3D token groups. Each group combines an instance token for entity-level identity with anchor tokens for local geometry and appearance, which are decoded into 3D Gaussians. Learned with differentiable rendering and joint reconstruction-segmentation supervision, the representation requires no 3D annotations. The authors report stronger class-agnostic instance segmentation than per-scene optimization baselines while remaining competitive for novel-view synthesis. The same groups support object removal, translation, insertion, and open-vocabulary 3D instance retrieval.
No heat snapshots are available in the last 24 hours.