InstruGen: Automatic Instruction Generation for Vision-and-Language Navigation Via Large Multimodal Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yan, Yu, Xu, Rongtao, Zhang, Jiazhao, Li, Peiyang, Liang, Xiaodan, Yin, Jianqin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910703049244672
author Yan, Yu
Xu, Rongtao
Zhang, Jiazhao
Li, Peiyang
Liang, Xiaodan
Yin, Jianqin
author_facet Yan, Yu
Xu, Rongtao
Zhang, Jiazhao
Li, Peiyang
Liang, Xiaodan
Yin, Jianqin
contents Recent research on Vision-and-Language Navigation (VLN) indicates that agents suffer from poor generalization in unseen environments due to the lack of realistic training environments and high-quality path-instruction pairs. Most existing methods for constructing realistic navigation scenes have high costs, and the extension of instructions mainly relies on predefined templates or rules, lacking adaptability. To alleviate the issue, we propose InstruGen, a VLN path-instruction pairs generation paradigm. Specifically, we use YouTube house tour videos as realistic navigation scenes and leverage the powerful visual understanding and generation abilities of large multimodal models (LMMs) to automatically generate diverse and high-quality VLN path-instruction pairs. Our method generates navigation instructions with different granularities and achieves fine-grained alignment between instructions and visual observations, which was difficult to achieve with previous methods. Additionally, we design a multi-stage verification mechanism to reduce hallucinations and inconsistency of LMMs. Experimental results demonstrate that agents trained with path-instruction pairs generated by InstruGen achieves state-of-the-art performance on the R2R and RxR benchmarks, particularly in unseen environments. Code is available at https://github.com/yanyu0526/InstruGen.
format Preprint
id arxiv_https___arxiv_org_abs_2411_11394
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle InstruGen: Automatic Instruction Generation for Vision-and-Language Navigation Via Large Multimodal Models
Yan, Yu
Xu, Rongtao
Zhang, Jiazhao
Li, Peiyang
Liang, Xiaodan
Yin, Jianqin
Robotics
Recent research on Vision-and-Language Navigation (VLN) indicates that agents suffer from poor generalization in unseen environments due to the lack of realistic training environments and high-quality path-instruction pairs. Most existing methods for constructing realistic navigation scenes have high costs, and the extension of instructions mainly relies on predefined templates or rules, lacking adaptability. To alleviate the issue, we propose InstruGen, a VLN path-instruction pairs generation paradigm. Specifically, we use YouTube house tour videos as realistic navigation scenes and leverage the powerful visual understanding and generation abilities of large multimodal models (LMMs) to automatically generate diverse and high-quality VLN path-instruction pairs. Our method generates navigation instructions with different granularities and achieves fine-grained alignment between instructions and visual observations, which was difficult to achieve with previous methods. Additionally, we design a multi-stage verification mechanism to reduce hallucinations and inconsistency of LMMs. Experimental results demonstrate that agents trained with path-instruction pairs generated by InstruGen achieves state-of-the-art performance on the R2R and RxR benchmarks, particularly in unseen environments. Code is available at https://github.com/yanyu0526/InstruGen.
title InstruGen: Automatic Instruction Generation for Vision-and-Language Navigation Via Large Multimodal Models
topic Robotics
url https://arxiv.org/abs/2411.11394