CATNAV: Cached Vision-Language Traversability for Efficient Zero-Shot Robot Navigation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Potnis, Aditya, Affonso, Francisco, Gummadi, Shreya, Uppalapati, Naveen Kumar, Chowdhary, Girish
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912980159954944
author Potnis, Aditya
Affonso, Francisco
Gummadi, Shreya
Uppalapati, Naveen Kumar
Chowdhary, Girish
author_facet Potnis, Aditya
Affonso, Francisco
Gummadi, Shreya
Uppalapati, Naveen Kumar
Chowdhary, Girish
contents Navigating unstructured environments requires assessing traversal risk relative to a robot's physical capabilities, a challenge that varies across embodiments. We present CATNAV, a cost-aware traversability navigation framework that leverages multimodal LLMs for zero-shot, embodiment-aware costmap generation without task-specific training. We introduce a visuosemantic caching mechanism that detects scene novelty and reuses prior risk assessments for semantically similar frames, reducing online VLM queries by 85.7%. Furthermore, we introduce a VLM-based trajectory selection module that evaluates proposals through visual reasoning to choose the safest path given behavioral constraints. We evaluate CATNAV on a quadruped robot across indoor and outdoor unstructured environments, comparing against state-of-the-art vision-language-action baselines. Across five navigation tasks, CATNAV achieves 10 percentage point higher average goal-reaching rate and 33% fewer behavioral constraint violations.
format Preprint
id arxiv_https___arxiv_org_abs_2603_22800
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CATNAV: Cached Vision-Language Traversability for Efficient Zero-Shot Robot Navigation
Potnis, Aditya
Affonso, Francisco
Gummadi, Shreya
Uppalapati, Naveen Kumar
Chowdhary, Girish
Robotics
Navigating unstructured environments requires assessing traversal risk relative to a robot's physical capabilities, a challenge that varies across embodiments. We present CATNAV, a cost-aware traversability navigation framework that leverages multimodal LLMs for zero-shot, embodiment-aware costmap generation without task-specific training. We introduce a visuosemantic caching mechanism that detects scene novelty and reuses prior risk assessments for semantically similar frames, reducing online VLM queries by 85.7%. Furthermore, we introduce a VLM-based trajectory selection module that evaluates proposals through visual reasoning to choose the safest path given behavioral constraints. We evaluate CATNAV on a quadruped robot across indoor and outdoor unstructured environments, comparing against state-of-the-art vision-language-action baselines. Across five navigation tasks, CATNAV achieves 10 percentage point higher average goal-reaching rate and 33% fewer behavioral constraint violations.
title CATNAV: Cached Vision-Language Traversability for Efficient Zero-Shot Robot Navigation
topic Robotics
url https://arxiv.org/abs/2603.22800