OmniNav: A Unified Framework for Prospective Exploration and Visual-Language Navigation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Xue, Xinda, Hu, Junjun, Luo, Minghua, Xie, Shichao, Chen, Jintao, Xie, Zixun, Quan, Kuichen, Guo, Wei, Xu, Mu, Chu, Zedong
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911357288316928
author Xue, Xinda
Hu, Junjun
Luo, Minghua
Xie, Shichao
Chen, Jintao
Xie, Zixun
Quan, Kuichen
Guo, Wei
Xu, Mu
Chu, Zedong
author_facet Xue, Xinda
Hu, Junjun
Luo, Minghua
Xie, Shichao
Chen, Jintao
Xie, Zixun
Quan, Kuichen
Guo, Wei
Xu, Mu
Chu, Zedong
contents Embodied navigation presents a core challenge for intelligent robots, requiring the comprehension of visual environments, natural language instructions, and autonomous exploration. Existing models often fall short in offering a unified solution across diverse navigation paradigms, resulting in low success rates and limited generalization. We introduce OmniNav, a unified framework addressing instruct-goal, object-goal, point-goal navigation, and frontier-based exploration within a single architecture. Our approach features a lightweight, low-latency policy that accurately predicts continuous-space waypoints (coordinates and orientations). This policy surpasses action-chunk methods in precision and supports real-world deployment at control frequencies up to 5 Hz. Architecturally, OmniNav employs a fast-slow system design: a fast module generates waypoints using short-horizon visual context and subtasks, while a slow module performs deliberative planning with long-horizon observations and candidate frontiers to select subsequent subgoals and subtasks. This collaboration enhances path efficiency and maintains trajectory coherence, particularly in exploration and memory-intensive scenarios. Crucially, we identify that the primary bottleneck isn't merely navigation policy learning, but a robust understanding of general instructions and objects. To boost generalization, OmniNav integrates large-scale, general-purpose training datasets, including those for image captioning and visual recognition, into a joint multi-task regimen. This significantly improves success rates and robustness. Extensive experiments confirm OmniNav's state-of-the-art performance across various navigation benchmarks, with real-world deployment further validating its efficacy. OmniNav provides practical insights for embodied navigation, charting a scalable path towards versatile, highly generalizable robotic intelligence.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25687
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OmniNav: A Unified Framework for Prospective Exploration and Visual-Language Navigation
Xue, Xinda
Hu, Junjun
Luo, Minghua
Xie, Shichao
Chen, Jintao
Xie, Zixun
Quan, Kuichen
Guo, Wei
Xu, Mu
Chu, Zedong
Robotics
Embodied navigation presents a core challenge for intelligent robots, requiring the comprehension of visual environments, natural language instructions, and autonomous exploration. Existing models often fall short in offering a unified solution across diverse navigation paradigms, resulting in low success rates and limited generalization. We introduce OmniNav, a unified framework addressing instruct-goal, object-goal, point-goal navigation, and frontier-based exploration within a single architecture. Our approach features a lightweight, low-latency policy that accurately predicts continuous-space waypoints (coordinates and orientations). This policy surpasses action-chunk methods in precision and supports real-world deployment at control frequencies up to 5 Hz. Architecturally, OmniNav employs a fast-slow system design: a fast module generates waypoints using short-horizon visual context and subtasks, while a slow module performs deliberative planning with long-horizon observations and candidate frontiers to select subsequent subgoals and subtasks. This collaboration enhances path efficiency and maintains trajectory coherence, particularly in exploration and memory-intensive scenarios. Crucially, we identify that the primary bottleneck isn't merely navigation policy learning, but a robust understanding of general instructions and objects. To boost generalization, OmniNav integrates large-scale, general-purpose training datasets, including those for image captioning and visual recognition, into a joint multi-task regimen. This significantly improves success rates and robustness. Extensive experiments confirm OmniNav's state-of-the-art performance across various navigation benchmarks, with real-world deployment further validating its efficacy. OmniNav provides practical insights for embodied navigation, charting a scalable path towards versatile, highly generalizable robotic intelligence.
title OmniNav: A Unified Framework for Prospective Exploration and Visual-Language Navigation
topic Robotics
url https://arxiv.org/abs/2509.25687