ABot-N1 proposes a slow-fast architecture for general visual-language navigation. A slow vision-language reasoner produces explicit linguistic reasoning and a pixel-space goal, while a fast action expert uses those signals to generate continuous waypoints at the native control frequency. The pixel anchors act as a shared interface across point-goal, object-goal, POI-goal, instruction-following, and person-following tasks. According to the supplied abstract, ABot-N1 improves POI arrival by 35.0 percentage points to 77.3%, reaches 95.4% and 92.9% success rates in complex indoor and outdoor scenes, and releases new Point-Goal and POI-Goal benchmarks.
No heat snapshots are available in the last 24 hours.