Navigating Beyond Instructions: Vision-and-Language Navigation in Obstructed Environments

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hong, Haodong, Wang, Sen, Huang, Zi, Wu, Qi, Liu, Jiajun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929444430544896
author Hong, Haodong
Wang, Sen
Huang, Zi
Wu, Qi
Liu, Jiajun
author_facet Hong, Haodong
Wang, Sen
Huang, Zi
Wu, Qi
Liu, Jiajun
contents Real-world navigation often involves dealing with unexpected obstructions such as closed doors, moved objects, and unpredictable entities. However, mainstream Vision-and-Language Navigation (VLN) tasks typically assume instructions perfectly align with the fixed and predefined navigation graphs without any obstructions. This assumption overlooks potential discrepancies in actual navigation graphs and given instructions, which can cause major failures for both indoor and outdoor agents. To address this issue, we integrate diverse obstructions into the R2R dataset by modifying both the navigation graphs and visual observations, introducing an innovative dataset and task, R2R with UNexpected Obstructions (R2R-UNO). R2R-UNO contains various types and numbers of path obstructions to generate instruction-reality mismatches for VLN research. Experiments on R2R-UNO reveal that state-of-the-art VLN methods inevitably encounter significant challenges when facing such mismatches, indicating that they rigidly follow instructions rather than navigate adaptively. Therefore, we propose a novel method called ObVLN (Obstructed VLN), which includes a curriculum training strategy and virtual graph construction to help agents effectively adapt to obstructed environments. Empirical results show that ObVLN not only maintains robust performance in unobstructed scenarios but also achieves a substantial performance advantage with unexpected obstructions.
format Preprint
id arxiv_https___arxiv_org_abs_2407_21452
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Navigating Beyond Instructions: Vision-and-Language Navigation in Obstructed Environments
Hong, Haodong
Wang, Sen
Huang, Zi
Wu, Qi
Liu, Jiajun
Robotics
Computation and Language
Computer Vision and Pattern Recognition
Real-world navigation often involves dealing with unexpected obstructions such as closed doors, moved objects, and unpredictable entities. However, mainstream Vision-and-Language Navigation (VLN) tasks typically assume instructions perfectly align with the fixed and predefined navigation graphs without any obstructions. This assumption overlooks potential discrepancies in actual navigation graphs and given instructions, which can cause major failures for both indoor and outdoor agents. To address this issue, we integrate diverse obstructions into the R2R dataset by modifying both the navigation graphs and visual observations, introducing an innovative dataset and task, R2R with UNexpected Obstructions (R2R-UNO). R2R-UNO contains various types and numbers of path obstructions to generate instruction-reality mismatches for VLN research. Experiments on R2R-UNO reveal that state-of-the-art VLN methods inevitably encounter significant challenges when facing such mismatches, indicating that they rigidly follow instructions rather than navigate adaptively. Therefore, we propose a novel method called ObVLN (Obstructed VLN), which includes a curriculum training strategy and virtual graph construction to help agents effectively adapt to obstructed environments. Empirical results show that ObVLN not only maintains robust performance in unobstructed scenarios but also achieves a substantial performance advantage with unexpected obstructions.
title Navigating Beyond Instructions: Vision-and-Language Navigation in Obstructed Environments
topic Robotics
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.21452