POINav: Benchmarking and Enhancing Final-Meters Arrival in Real-World Vision-Language Navigation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916053847638016 |
|---|---|
| author | Gong, Ruiyan Zhang, Meisheng Zhao, Yuxiang Sun, Mingchao Shen, Yanfen Chu, Zedong Gu, Zhining Guo, Wei Cheng, Xiaolong Li, Qiming Niu, Kangning Zhu, Yanqing Wu, Xiaolong Li, Tianlun Xu, Mu |
| author_facet | Gong, Ruiyan Zhang, Meisheng Zhao, Yuxiang Sun, Mingchao Shen, Yanfen Chu, Zedong Gu, Zhining Guo, Wei Cheng, Xiaolong Li, Qiming Niu, Kangning Zhu, Yanqing Wu, Xiaolong Li, Tianlun Xu, Mu |
| contents | Real-world navigation is fundamentally driven by Points of Interest (POIs), yet reaching a precise POI remains a critical "final-meters" challenge. Existing Vision-Language Navigation (VLN) benchmarks of POI-goal navigation often suffer from coarse granularity or significant sim-to-real gaps due to generated scene. To bridge this gap, we present POINav-Bench, the first benchmark designed for closed-loop evaluation of real-world POI-goal navigation. It comprises 11 commercial areas reconstructed from real-world captures using 3D Gaussian Splatting (3DGS), covering 126,398 $m^{2}$ in total and spanning 163 distinct POIs. With traversability-aware annotations and reference trajectories, POINav-Bench enables high-fidelity evaluation of navigation agents in realistic, POI-rich real-world environments. Building on this, we propose the POINav Brain-Action Framework where a Brain module performs POI-grounded reasoning to guide an Action module in predicting continuous waypoints for real-world execution. We further curate the POINav-Dataset, containing 70K real-world signage-entrance pairs. Experiments show that our framework provides a viable path toward refining real-world POI-goal navigation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_28237 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | POINav: Benchmarking and Enhancing Final-Meters Arrival in Real-World Vision-Language Navigation Gong, Ruiyan Zhang, Meisheng Zhao, Yuxiang Sun, Mingchao Shen, Yanfen Chu, Zedong Gu, Zhining Guo, Wei Cheng, Xiaolong Li, Qiming Niu, Kangning Zhu, Yanqing Wu, Xiaolong Li, Tianlun Xu, Mu Robotics Computer Vision and Pattern Recognition Real-world navigation is fundamentally driven by Points of Interest (POIs), yet reaching a precise POI remains a critical "final-meters" challenge. Existing Vision-Language Navigation (VLN) benchmarks of POI-goal navigation often suffer from coarse granularity or significant sim-to-real gaps due to generated scene. To bridge this gap, we present POINav-Bench, the first benchmark designed for closed-loop evaluation of real-world POI-goal navigation. It comprises 11 commercial areas reconstructed from real-world captures using 3D Gaussian Splatting (3DGS), covering 126,398 $m^{2}$ in total and spanning 163 distinct POIs. With traversability-aware annotations and reference trajectories, POINav-Bench enables high-fidelity evaluation of navigation agents in realistic, POI-rich real-world environments. Building on this, we propose the POINav Brain-Action Framework where a Brain module performs POI-grounded reasoning to guide an Action module in predicting continuous waypoints for real-world execution. We further curate the POINav-Dataset, containing 70K real-world signage-entrance pairs. Experiments show that our framework provides a viable path toward refining real-world POI-goal navigation. |
| title | POINav: Benchmarking and Enhancing Final-Meters Arrival in Real-World Vision-Language Navigation |
| topic | Robotics Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2605.28237 |