Robostral Navigate is an 8B vision-language navigation model that uses only streams of monocular RGB images and predicts waypoints by pointing to the next target location in the current view. The authors generate 2.4 million trajectories across 350,000 simulated scenes, and combine prefix-caching, a tree-based attention mask, and reinforcement learning. The reported recipe reduces training tokens by 22x and cuts training time from months to days. It achieves 77.4% success on R2R-CE and 75.1% on RxR-CE, outperforming reported monocular baselines.
No heat snapshots are available in the last 24 hours.