The paper introduces VTM-Nav, a training-free ObjectNav framework that preserves scene-scoped experience across independently initialized episodes. Its Hierarchical Visual-Topological Memory organizes visual records by coarse room topology, separates in-room evidence from remotely visible evidence, and stores successful approach cues. During each request, the agent re-localizes itself, retrieves target-relevant records from plausible rooms, grounds guidance in candidates from the current observation, and uses a conservative execution guard for local failures. Under matched 40-step evaluations, VTM-Nav improves success rate over the memory-reset WMNav baseline by 4.6, 2.0, and 0.8 points on HM3D v0.1, HM3D v0.2, and MP3D, respectively.
No heat snapshots are available in the last 24 hours.