General-Purpose Robotic Navigation via LVLM-Orchestrated Perception, Reasoning, and Acting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lange, Bernard, Yildiz, Anil, Arief, Mansur, Khattak, Shehryar, Kochenderfer, Mykel, Georgakis, Georgios
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911215733702656
author Lange, Bernard
Yildiz, Anil
Arief, Mansur
Khattak, Shehryar
Kochenderfer, Mykel
Georgakis, Georgios
author_facet Lange, Bernard
Yildiz, Anil
Arief, Mansur
Khattak, Shehryar
Kochenderfer, Mykel
Georgakis, Georgios
contents Developing general-purpose navigation policies for unknown environments remains a core challenge in robotics. Most existing systems rely on task-specific neural networks and fixed information flows, limiting their generalizability. Large Vision-Language Models (LVLMs) offer a promising alternative by embedding human-like knowledge for reasoning and planning, but prior LVLM-robot integrations have largely depended on pre-mapped spaces, hard-coded representations, and rigid control logic. We introduce the Agentic Robotic Navigation Architecture (ARNA), a general-purpose framework that equips an LVLM-based agent with a library of perception, reasoning, and navigation tools drawn from modern robotic stacks. At runtime, the agent autonomously defines and executes task-specific workflows that iteratively query modules, reason over multimodal inputs, and select navigation actions. This agentic formulation enables robust navigation and reasoning in previously unmapped environments, offering a new perspective on robotic stack design. Evaluated in Habitat Lab on the HM-EQA benchmark, ARNA outperforms state-of-the-art EQA-specific approaches. Qualitative results on RxR and custom tasks further demonstrate its ability to generalize across a broad range of navigation challenges.
format Preprint
id arxiv_https___arxiv_org_abs_2506_17462
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle General-Purpose Robotic Navigation via LVLM-Orchestrated Perception, Reasoning, and Acting
Lange, Bernard
Yildiz, Anil
Arief, Mansur
Khattak, Shehryar
Kochenderfer, Mykel
Georgakis, Georgios
Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Developing general-purpose navigation policies for unknown environments remains a core challenge in robotics. Most existing systems rely on task-specific neural networks and fixed information flows, limiting their generalizability. Large Vision-Language Models (LVLMs) offer a promising alternative by embedding human-like knowledge for reasoning and planning, but prior LVLM-robot integrations have largely depended on pre-mapped spaces, hard-coded representations, and rigid control logic. We introduce the Agentic Robotic Navigation Architecture (ARNA), a general-purpose framework that equips an LVLM-based agent with a library of perception, reasoning, and navigation tools drawn from modern robotic stacks. At runtime, the agent autonomously defines and executes task-specific workflows that iteratively query modules, reason over multimodal inputs, and select navigation actions. This agentic formulation enables robust navigation and reasoning in previously unmapped environments, offering a new perspective on robotic stack design. Evaluated in Habitat Lab on the HM-EQA benchmark, ARNA outperforms state-of-the-art EQA-specific approaches. Qualitative results on RxR and custom tasks further demonstrate its ability to generalize across a broad range of navigation challenges.
title General-Purpose Robotic Navigation via LVLM-Orchestrated Perception, Reasoning, and Acting
topic Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.17462