CATNAV: Cached Vision-Language Traversability for Efficient Zero-Shot Robot Navigation
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866912980159954944 |
|---|---|
| author | Potnis, Aditya Affonso, Francisco Gummadi, Shreya Uppalapati, Naveen Kumar Chowdhary, Girish |
| author_facet | Potnis, Aditya Affonso, Francisco Gummadi, Shreya Uppalapati, Naveen Kumar Chowdhary, Girish |
| contents | Navigating unstructured environments requires assessing traversal risk relative to a robot's physical capabilities, a challenge that varies across embodiments. We present CATNAV, a cost-aware traversability navigation framework that leverages multimodal LLMs for zero-shot, embodiment-aware costmap generation without task-specific training. We introduce a visuosemantic caching mechanism that detects scene novelty and reuses prior risk assessments for semantically similar frames, reducing online VLM queries by 85.7%. Furthermore, we introduce a VLM-based trajectory selection module that evaluates proposals through visual reasoning to choose the safest path given behavioral constraints. We evaluate CATNAV on a quadruped robot across indoor and outdoor unstructured environments, comparing against state-of-the-art vision-language-action baselines. Across five navigation tasks, CATNAV achieves 10 percentage point higher average goal-reaching rate and 33% fewer behavioral constraint violations. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_22800 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | CATNAV: Cached Vision-Language Traversability for Efficient Zero-Shot Robot Navigation Potnis, Aditya Affonso, Francisco Gummadi, Shreya Uppalapati, Naveen Kumar Chowdhary, Girish Robotics Navigating unstructured environments requires assessing traversal risk relative to a robot's physical capabilities, a challenge that varies across embodiments. We present CATNAV, a cost-aware traversability navigation framework that leverages multimodal LLMs for zero-shot, embodiment-aware costmap generation without task-specific training. We introduce a visuosemantic caching mechanism that detects scene novelty and reuses prior risk assessments for semantically similar frames, reducing online VLM queries by 85.7%. Furthermore, we introduce a VLM-based trajectory selection module that evaluates proposals through visual reasoning to choose the safest path given behavioral constraints. We evaluate CATNAV on a quadruped robot across indoor and outdoor unstructured environments, comparing against state-of-the-art vision-language-action baselines. Across five navigation tasks, CATNAV achieves 10 percentage point higher average goal-reaching rate and 33% fewer behavioral constraint violations. |
| title | CATNAV: Cached Vision-Language Traversability for Efficient Zero-Shot Robot Navigation |
| topic | Robotics |
| url | https://arxiv.org/abs/2603.22800 |