PathPainter: Transferring the Generalization Ability of Image Generation Models to Embodied Navigation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866910200868372480 |
|---|---|
| author | Wang, Yijin Tian, Yuru Huang, Xijie Gai, Weiqi Zhu, Mo Zhou, Xin Wu, Yuze Gao, Fei |
| author_facet | Wang, Yijin Tian, Yuru Huang, Xijie Gai, Weiqi Zhu, Mo Zhou, Xin Wu, Yuze Gao, Fei |
| contents | Bird's-eye-view (BEV) images have been widely demonstrated to provide valuable prior information for navigation. Given the global information provided by such views, two key challenges remain: how to fully exploit this information and how to reliably use it during execution. In this paper, we propose a navigation system that uses BEV images as global priors and is designed for ground and near-ground robotic platforms. The system employs an image generation model to interpret human intent from natural language, identify the target destination, and generate traversability masks. During execution, we introduce cross-view localization to align the robot's odometry with the BEV map and mitigate long-term drift in conventional odometry. We conduct extensive benchmark experiments to evaluate the proposed method and further validate it on a UAV platform. Using only a conventional local motion planner, the UAV successfully completes a 160-meter outdoor long-range navigation task. This work demonstrates how the world-understanding capabilities of foundation models can be transferred to embodied navigation, enabling robots to benefit from the strong generalization ability of existing image generation models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_07496 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | PathPainter: Transferring the Generalization Ability of Image Generation Models to Embodied Navigation Wang, Yijin Tian, Yuru Huang, Xijie Gai, Weiqi Zhu, Mo Zhou, Xin Wu, Yuze Gao, Fei Robotics Bird's-eye-view (BEV) images have been widely demonstrated to provide valuable prior information for navigation. Given the global information provided by such views, two key challenges remain: how to fully exploit this information and how to reliably use it during execution. In this paper, we propose a navigation system that uses BEV images as global priors and is designed for ground and near-ground robotic platforms. The system employs an image generation model to interpret human intent from natural language, identify the target destination, and generate traversability masks. During execution, we introduce cross-view localization to align the robot's odometry with the BEV map and mitigate long-term drift in conventional odometry. We conduct extensive benchmark experiments to evaluate the proposed method and further validate it on a UAV platform. Using only a conventional local motion planner, the UAV successfully completes a 160-meter outdoor long-range navigation task. This work demonstrates how the world-understanding capabilities of foundation models can be transferred to embodied navigation, enabling robots to benefit from the strong generalization ability of existing image generation models. |
| title | PathPainter: Transferring the Generalization Ability of Image Generation Models to Embodied Navigation |
| topic | Robotics |
| url | https://arxiv.org/abs/2605.07496 |