General-Purpose Robotic Navigation via LVLM-Orchestrated Perception, Reasoning, and Acting
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911215733702656 |
|---|---|
| author | Lange, Bernard Yildiz, Anil Arief, Mansur Khattak, Shehryar Kochenderfer, Mykel Georgakis, Georgios |
| author_facet | Lange, Bernard Yildiz, Anil Arief, Mansur Khattak, Shehryar Kochenderfer, Mykel Georgakis, Georgios |
| contents | Developing general-purpose navigation policies for unknown environments remains a core challenge in robotics. Most existing systems rely on task-specific neural networks and fixed information flows, limiting their generalizability. Large Vision-Language Models (LVLMs) offer a promising alternative by embedding human-like knowledge for reasoning and planning, but prior LVLM-robot integrations have largely depended on pre-mapped spaces, hard-coded representations, and rigid control logic. We introduce the Agentic Robotic Navigation Architecture (ARNA), a general-purpose framework that equips an LVLM-based agent with a library of perception, reasoning, and navigation tools drawn from modern robotic stacks. At runtime, the agent autonomously defines and executes task-specific workflows that iteratively query modules, reason over multimodal inputs, and select navigation actions. This agentic formulation enables robust navigation and reasoning in previously unmapped environments, offering a new perspective on robotic stack design. Evaluated in Habitat Lab on the HM-EQA benchmark, ARNA outperforms state-of-the-art EQA-specific approaches. Qualitative results on RxR and custom tasks further demonstrate its ability to generalize across a broad range of navigation challenges. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_17462 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | General-Purpose Robotic Navigation via LVLM-Orchestrated Perception, Reasoning, and Acting Lange, Bernard Yildiz, Anil Arief, Mansur Khattak, Shehryar Kochenderfer, Mykel Georgakis, Georgios Robotics Artificial Intelligence Computer Vision and Pattern Recognition Developing general-purpose navigation policies for unknown environments remains a core challenge in robotics. Most existing systems rely on task-specific neural networks and fixed information flows, limiting their generalizability. Large Vision-Language Models (LVLMs) offer a promising alternative by embedding human-like knowledge for reasoning and planning, but prior LVLM-robot integrations have largely depended on pre-mapped spaces, hard-coded representations, and rigid control logic. We introduce the Agentic Robotic Navigation Architecture (ARNA), a general-purpose framework that equips an LVLM-based agent with a library of perception, reasoning, and navigation tools drawn from modern robotic stacks. At runtime, the agent autonomously defines and executes task-specific workflows that iteratively query modules, reason over multimodal inputs, and select navigation actions. This agentic formulation enables robust navigation and reasoning in previously unmapped environments, offering a new perspective on robotic stack design. Evaluated in Habitat Lab on the HM-EQA benchmark, ARNA outperforms state-of-the-art EQA-specific approaches. Qualitative results on RxR and custom tasks further demonstrate its ability to generalize across a broad range of navigation challenges. |
| title | General-Purpose Robotic Navigation via LVLM-Orchestrated Perception, Reasoning, and Acting |
| topic | Robotics Artificial Intelligence Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2506.17462 |