The paper introduces Future-State-Conditioned VLN (FSC-VLN), a deployable vision-language navigation policy trained with privileged future-state supervision. During training, a future-query representation is aligned with a frozen visual embedding from an expert trajectory Δ steps ahead, distilling predictive information into the policy state. At inference, the model uses only past and current observations and adds two learned prefix tokens to the baseline pattern. On R2R val-unseen, FSC-VLN reportedly improves success rate (SR), oracle success rate (OSR), and success weighted by path length (SPL) over a StreamVLN-style baseline under two training-data regimes, with larger gains on long-horizon episodes.
No heat snapshots are available in the last 24 hours.