This paper presents a generalized data-association-free object SLAM framework that jointly estimates robot poses, landmark positions, landmark semantics, and latent measurement associations. It combines positional observations with either object-class labels or real-valued feature vectors from visual foundation models. A semi-incremental estimation scheme is introduced to balance accuracy and computational efficiency, while the paper also provides principles and heuristics for estimating the number of landmarks. According to the abstract, experiments on synthetic and real-world datasets show improvements over strong baselines, although detailed datasets, metrics, and numerical results are not available in the supplied summary.
No heat snapshots are available in the last 24 hours.