Endowing Embodied Agents with Spatial Reasoning Capabilities for Vision-and-Language Navigation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Bai, Qianqian, Chen, Zhongpu, Luo, Ling, Du, Huaming, Lei, Yuqian, Jiao, Ziyun
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908859454455808
author Bai, Qianqian
Chen, Zhongpu
Luo, Ling
Du, Huaming
Lei, Yuqian
Jiao, Ziyun
author_facet Bai, Qianqian
Chen, Zhongpu
Luo, Ling
Du, Huaming
Lei, Yuqian
Jiao, Ziyun
contents Enhancing the spatial perception capabilities of mobile robots is crucial for achieving embodied Vision-and-Language Navigation (VLN). Although significant progress has been made in simulated environments, directly transferring these capabilities to real-world scenarios often results in severe hallucination phenomena, causing robots to lose effective spatial awareness. To address this issue, we propose BrainNav, a bio-inspired spatial cognitive navigation framework inspired by biological spatial cognition theories and cognitive map theory. BrainNav integrates dual-map (coordinate map and topological map) and dual-orientation (relative orientation and absolute orientation) strategies, enabling real-time navigation through dynamic scene capture and path planning. Its five core modules-Hippocampal Memory Hub, Visual Cortex Perception Engine, Parietal Spatial Constructor, Prefrontal Decision Center, and Cerebellar Motion Execution Unit-mimic biological cognitive functions to reduce spatial hallucinations and enhance adaptability. Validated in a zero-shot real-world lab environment using the Limo Pro robot, BrainNav, compatible with GPT-4, outperforms existing State-of-the-Art (SOTA) Vision-and-Language Navigation in Continuous Environments (VLN-CE) methods without fine-tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2504_08806
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Endowing Embodied Agents with Spatial Reasoning Capabilities for Vision-and-Language Navigation
Bai, Qianqian
Chen, Zhongpu
Luo, Ling
Du, Huaming
Lei, Yuqian
Jiao, Ziyun
Artificial Intelligence
Robotics
Enhancing the spatial perception capabilities of mobile robots is crucial for achieving embodied Vision-and-Language Navigation (VLN). Although significant progress has been made in simulated environments, directly transferring these capabilities to real-world scenarios often results in severe hallucination phenomena, causing robots to lose effective spatial awareness. To address this issue, we propose BrainNav, a bio-inspired spatial cognitive navigation framework inspired by biological spatial cognition theories and cognitive map theory. BrainNav integrates dual-map (coordinate map and topological map) and dual-orientation (relative orientation and absolute orientation) strategies, enabling real-time navigation through dynamic scene capture and path planning. Its five core modules-Hippocampal Memory Hub, Visual Cortex Perception Engine, Parietal Spatial Constructor, Prefrontal Decision Center, and Cerebellar Motion Execution Unit-mimic biological cognitive functions to reduce spatial hallucinations and enhance adaptability. Validated in a zero-shot real-world lab environment using the Limo Pro robot, BrainNav, compatible with GPT-4, outperforms existing State-of-the-Art (SOTA) Vision-and-Language Navigation in Continuous Environments (VLN-CE) methods without fine-tuning.
title Endowing Embodied Agents with Spatial Reasoning Capabilities for Vision-and-Language Navigation
topic Artificial Intelligence
Robotics
url https://arxiv.org/abs/2504.08806