Learning Goal-Oriented Vision-and-Language Navigation with Self-Improving Demonstrations at Scale

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Songze, Wang, Zun, Zhou, Gengze, Li, Jialu, Zeng, Xiangyu, Gong, Ziyang, Wang, Limin, Qiao, Yu, Wu, Qi, Bansal, Mohit, Wang, Yi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914403557834752
author Li, Songze
Wang, Zun
Zhou, Gengze
Li, Jialu
Zeng, Xiangyu
Gong, Ziyang
Wang, Limin
Qiao, Yu
Wu, Qi
Bansal, Mohit
Wang, Yi
author_facet Li, Songze
Wang, Zun
Zhou, Gengze
Li, Jialu
Zeng, Xiangyu
Gong, Ziyang
Wang, Limin
Qiao, Yu
Wu, Qi
Bansal, Mohit
Wang, Yi
contents Goal-oriented vision-language navigation requires robust exploration capabilities for agents to navigate to specified goals in unknown environments without step-by-step instructions. Existing methods tend to exclusively utilize shortest-path trajectories, lacking effective exploration priors for training navigation agents. To address the above challenges, we present SID, a goal-oriented vision-and-language navigation learning approach with Self-Improving Demonstrations. Specifically, SID learns an initial agent on the shortest-path data sampled from environments and then leverages this agent to generate novel exploration trajectories. The novel rollouts provide demonstrations with stronger exploration strategies to train a better agent, which in turn produces higher-quality agent demonstrations for the next round of training. We show that this iterative self-improving pipeline readily scales to new environments, and the resulting demonstrations are highly transferable, elevating the performance ceiling across a variety of vision-and-language navigation tasks. Extensive experiments demonstrate that SID significantly boosts the exploration capabilities and generalization of navigation agents. The resulting agent achieves new state-of-the-art performance on goal-oriented vision-and-language navigation benchmarks, including REVERIE, SOON as well as strong transferability to object-goal navigation and VLN-CE. It notably achieves a 50.9% success rate on the unseen validation splits of SOON, surpassing prior leading approaches by a margin of 13.9%.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24910
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning Goal-Oriented Vision-and-Language Navigation with Self-Improving Demonstrations at Scale
Li, Songze
Wang, Zun
Zhou, Gengze
Li, Jialu
Zeng, Xiangyu
Gong, Ziyang
Wang, Limin
Qiao, Yu
Wu, Qi
Bansal, Mohit
Wang, Yi
Computer Vision and Pattern Recognition
Goal-oriented vision-language navigation requires robust exploration capabilities for agents to navigate to specified goals in unknown environments without step-by-step instructions. Existing methods tend to exclusively utilize shortest-path trajectories, lacking effective exploration priors for training navigation agents. To address the above challenges, we present SID, a goal-oriented vision-and-language navigation learning approach with Self-Improving Demonstrations. Specifically, SID learns an initial agent on the shortest-path data sampled from environments and then leverages this agent to generate novel exploration trajectories. The novel rollouts provide demonstrations with stronger exploration strategies to train a better agent, which in turn produces higher-quality agent demonstrations for the next round of training. We show that this iterative self-improving pipeline readily scales to new environments, and the resulting demonstrations are highly transferable, elevating the performance ceiling across a variety of vision-and-language navigation tasks. Extensive experiments demonstrate that SID significantly boosts the exploration capabilities and generalization of navigation agents. The resulting agent achieves new state-of-the-art performance on goal-oriented vision-and-language navigation benchmarks, including REVERIE, SOON as well as strong transferability to object-goal navigation and VLN-CE. It notably achieves a 50.9% success rate on the unseen validation splits of SOON, surpassing prior leading approaches by a margin of 13.9%.
title Learning Goal-Oriented Vision-and-Language Navigation with Self-Improving Demonstrations at Scale
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.24910