FlexVLN: Flexible Adaptation for Diverse Vision-and-Language Navigation Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Siqi, Qiao, Yanyuan, Wang, Qunbo, Guo, Longteng, Wei, Zhihua, Liu, Jing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912280343478272
author Zhang, Siqi
Qiao, Yanyuan
Wang, Qunbo
Guo, Longteng
Wei, Zhihua
Liu, Jing
author_facet Zhang, Siqi
Qiao, Yanyuan
Wang, Qunbo
Guo, Longteng
Wei, Zhihua
Liu, Jing
contents The aspiration of the Vision-and-Language Navigation (VLN) task has long been to develop an embodied agent with robust adaptability, capable of seamlessly transferring its navigation capabilities across various tasks. Despite remarkable advancements in recent years, most methods necessitate dataset-specific training, thereby lacking the capability to generalize across diverse datasets encompassing distinct types of instructions. Large language models (LLMs) have demonstrated exceptional reasoning and generalization abilities, exhibiting immense potential in robot action planning. In this paper, we propose FlexVLN, an innovative hierarchical approach to VLN that integrates the fundamental navigation ability of a supervised-learning-based Instruction Follower with the robust generalization ability of the LLM Planner, enabling effective generalization across diverse VLN datasets. Moreover, a verification mechanism and a multi-model integration mechanism are proposed to mitigate potential hallucinations by the LLM Planner and enhance execution accuracy of the Instruction Follower. We take REVERIE, SOON, and CVDN-target as out-of-domain datasets for assessing generalization ability. The generalization performance of FlexVLN surpasses that of all the previous methods to a large extent.
format Preprint
id arxiv_https___arxiv_org_abs_2503_13966
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FlexVLN: Flexible Adaptation for Diverse Vision-and-Language Navigation Tasks
Zhang, Siqi
Qiao, Yanyuan
Wang, Qunbo
Guo, Longteng
Wei, Zhihua
Liu, Jing
Computer Vision and Pattern Recognition
Robotics
The aspiration of the Vision-and-Language Navigation (VLN) task has long been to develop an embodied agent with robust adaptability, capable of seamlessly transferring its navigation capabilities across various tasks. Despite remarkable advancements in recent years, most methods necessitate dataset-specific training, thereby lacking the capability to generalize across diverse datasets encompassing distinct types of instructions. Large language models (LLMs) have demonstrated exceptional reasoning and generalization abilities, exhibiting immense potential in robot action planning. In this paper, we propose FlexVLN, an innovative hierarchical approach to VLN that integrates the fundamental navigation ability of a supervised-learning-based Instruction Follower with the robust generalization ability of the LLM Planner, enabling effective generalization across diverse VLN datasets. Moreover, a verification mechanism and a multi-model integration mechanism are proposed to mitigate potential hallucinations by the LLM Planner and enhance execution accuracy of the Instruction Follower. We take REVERIE, SOON, and CVDN-target as out-of-domain datasets for assessing generalization ability. The generalization performance of FlexVLN surpasses that of all the previous methods to a large extent.
title FlexVLN: Flexible Adaptation for Diverse Vision-and-Language Navigation Tasks
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2503.13966