AdaVLN: Towards Visual Language Navigation in Continuous Indoor Environments with Moving Humans

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Loh, Dillon, Bednarz, Tomasz, Xia, Xinxing, Guan, Frank
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911355951382528
author Loh, Dillon
Bednarz, Tomasz
Xia, Xinxing
Guan, Frank
author_facet Loh, Dillon
Bednarz, Tomasz
Xia, Xinxing
Guan, Frank
contents Visual Language Navigation is a task that challenges robots to navigate in realistic environments based on natural language instructions. While previous research has largely focused on static settings, real-world navigation must often contend with dynamic human obstacles. Hence, we propose an extension to the task, termed Adaptive Visual Language Navigation (AdaVLN), which seeks to narrow this gap. AdaVLN requires robots to navigate complex 3D indoor environments populated with dynamically moving human obstacles, adding a layer of complexity to navigation tasks that mimic the real-world. To support exploration of this task, we also present AdaVLN simulator and AdaR2R datasets. The AdaVLN simulator enables easy inclusion of fully animated human models directly into common datasets like Matterport3D. We also introduce a "freeze-time" mechanism for both the navigation task and simulator, which pauses world state updates during agent inference, enabling fair comparisons and experimental reproducibility across different hardware. We evaluate several baseline models on this task, analyze the unique challenges introduced by AdaVLN, and demonstrate its potential to bridge the sim-to-real gap in VLN research.
format Preprint
id arxiv_https___arxiv_org_abs_2411_18539
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AdaVLN: Towards Visual Language Navigation in Continuous Indoor Environments with Moving Humans
Loh, Dillon
Bednarz, Tomasz
Xia, Xinxing
Guan, Frank
Computer Vision and Pattern Recognition
Robotics
Visual Language Navigation is a task that challenges robots to navigate in realistic environments based on natural language instructions. While previous research has largely focused on static settings, real-world navigation must often contend with dynamic human obstacles. Hence, we propose an extension to the task, termed Adaptive Visual Language Navigation (AdaVLN), which seeks to narrow this gap. AdaVLN requires robots to navigate complex 3D indoor environments populated with dynamically moving human obstacles, adding a layer of complexity to navigation tasks that mimic the real-world. To support exploration of this task, we also present AdaVLN simulator and AdaR2R datasets. The AdaVLN simulator enables easy inclusion of fully animated human models directly into common datasets like Matterport3D. We also introduce a "freeze-time" mechanism for both the navigation task and simulator, which pauses world state updates during agent inference, enabling fair comparisons and experimental reproducibility across different hardware. We evaluate several baseline models on this task, analyze the unique challenges introduced by AdaVLN, and demonstrate its potential to bridge the sim-to-real gap in VLN research.
title AdaVLN: Towards Visual Language Navigation in Continuous Indoor Environments with Moving Humans
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2411.18539