The paper introduces Phi-Nav, an on-policy training framework for Vision-Language Navigation that addresses the semantic mismatch between exploratory trajectories and the original instructions. It uses a three-stage dual-supervision cycle: oracle-guided exploration, generation of a path-level hindsight instruction from collected visual observations, and a second imitation pass over the synthesized trajectory-instruction pair. On R2R-CE and RxR-CE, the authors report competitive performance while using only a fraction of the expert demonstrations required by current baselines. The abstract does not provide exact quantitative results.
No heat snapshots are available in the last 24 hours.