UnitedVLN: Generalizable Gaussian Splatting for Continuous Vision-Language Navigation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dai, Guangzhao, Zhao, Jian, Chen, Yuantao, Qin, Yusen, Zhao, Hao, Xie, Guosen, Yao, Yazhou, Shu, Xiangbo, Li, Xuelong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917957619154944
author Dai, Guangzhao
Zhao, Jian
Chen, Yuantao
Qin, Yusen
Zhao, Hao
Xie, Guosen
Yao, Yazhou
Shu, Xiangbo
Li, Xuelong
author_facet Dai, Guangzhao
Zhao, Jian
Chen, Yuantao
Qin, Yusen
Zhao, Hao
Xie, Guosen
Yao, Yazhou
Shu, Xiangbo
Li, Xuelong
contents Vision-and-Language Navigation (VLN), where an agent follows instructions to reach a target destination, has recently seen significant advancements. In contrast to navigation in discrete environments with predefined trajectories, VLN in Continuous Environments (VLN-CE) presents greater challenges, as the agent is free to navigate any unobstructed location and is more vulnerable to visual occlusions or blind spots. Recent approaches have attempted to address this by imagining future environments, either through predicted future visual images or semantic features, rather than relying solely on current observations. However, these RGB-based and feature-based methods lack intuitive appearance-level information or high-level semantic complexity crucial for effective navigation. To overcome these limitations, we introduce a novel, generalizable 3DGS-based pre-training paradigm, called UnitedVLN, which enables agents to better explore future environments by unitedly rendering high-fidelity 360 visual images and semantic features. UnitedVLN employs two key schemes: search-then-query sampling and separate-then-united rendering, which facilitate efficient exploitation of neural primitives, helping to integrate both appearance and semantic information for more robust navigation. Extensive experiments demonstrate that UnitedVLN outperforms state-of-the-art methods on existing VLN-CE benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2411_16053
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle UnitedVLN: Generalizable Gaussian Splatting for Continuous Vision-Language Navigation
Dai, Guangzhao
Zhao, Jian
Chen, Yuantao
Qin, Yusen
Zhao, Hao
Xie, Guosen
Yao, Yazhou
Shu, Xiangbo
Li, Xuelong
Computer Vision and Pattern Recognition
Artificial Intelligence
Vision-and-Language Navigation (VLN), where an agent follows instructions to reach a target destination, has recently seen significant advancements. In contrast to navigation in discrete environments with predefined trajectories, VLN in Continuous Environments (VLN-CE) presents greater challenges, as the agent is free to navigate any unobstructed location and is more vulnerable to visual occlusions or blind spots. Recent approaches have attempted to address this by imagining future environments, either through predicted future visual images or semantic features, rather than relying solely on current observations. However, these RGB-based and feature-based methods lack intuitive appearance-level information or high-level semantic complexity crucial for effective navigation. To overcome these limitations, we introduce a novel, generalizable 3DGS-based pre-training paradigm, called UnitedVLN, which enables agents to better explore future environments by unitedly rendering high-fidelity 360 visual images and semantic features. UnitedVLN employs two key schemes: search-then-query sampling and separate-then-united rendering, which facilitate efficient exploitation of neural primitives, helping to integrate both appearance and semantic information for more robust navigation. Extensive experiments demonstrate that UnitedVLN outperforms state-of-the-art methods on existing VLN-CE benchmarks.
title UnitedVLN: Generalizable Gaussian Splatting for Continuous Vision-Language Navigation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2411.16053