CapNav: Benchmarking Vision Language Models on Capability-conditioned Indoor Navigation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Su, Xia, Chen, Ruiqi, Liu, Benlin, Ma, Jingwei, Di, Zonglin, Krishna, Ranjay, Froehlich, Jon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VL-Nav: A Neuro-Symbolic Approach for Reasoning-based Vision-Language Navigation
von: Du, Yi, et al.
Veröffentlicht: (2025)
von: Du, Yi, et al.
Veröffentlicht: (2025)
NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models
von: Zhou, Gengze, et al.
Veröffentlicht: (2024)
von: Zhou, Gengze, et al.
Veröffentlicht: (2024)
SignNav: Leveraging Signage for Semantic Visual Navigation in Large-Scale Indoor Environments
von: Sun, Jian, et al.
Veröffentlicht: (2026)
von: Sun, Jian, et al.
Veröffentlicht: (2026)
NavQ: Learning a Q-Model for Foresighted Vision-and-Language Navigation
von: Xu, Peiran, et al.
Veröffentlicht: (2025)
von: Xu, Peiran, et al.
Veröffentlicht: (2025)
The One RING: a Robotic Indoor Navigation Generalist
von: Eftekhar, Ainaz, et al.
Veröffentlicht: (2024)
von: Eftekhar, Ainaz, et al.
Veröffentlicht: (2024)
CoNav: A Benchmark for Human-Centered Collaborative Navigation
von: Li, Changhao, et al.
Veröffentlicht: (2024)
von: Li, Changhao, et al.
Veröffentlicht: (2024)
VISTAv2: World Imagination for Indoor Vision-and-Language Navigation
von: Huang, Yanjia, et al.
Veröffentlicht: (2025)
von: Huang, Yanjia, et al.
Veröffentlicht: (2025)
VLS: Steering Pretrained Robot Policies via Vision-Language Models
von: Liu, Shuo, et al.
Veröffentlicht: (2026)
von: Liu, Shuo, et al.
Veröffentlicht: (2026)
Open-Nav: Exploring Zero-Shot Vision-and-Language Navigation in Continuous Environment with Open-Source LLMs
von: Qiao, Yanyuan, et al.
Veröffentlicht: (2024)
von: Qiao, Yanyuan, et al.
Veröffentlicht: (2024)
HaltNav: Reactive Visual Halting over Lightweight Topological Priors for Robust Vision-Language Navigation
von: Yu, Zihui, et al.
Veröffentlicht: (2026)
von: Yu, Zihui, et al.
Veröffentlicht: (2026)
Three-Step Nav: A Hierarchical Global-Local Planner for Zero-Shot Vision-and-Language Navigation
von: Zheng, Wanrong, et al.
Veröffentlicht: (2026)
von: Zheng, Wanrong, et al.
Veröffentlicht: (2026)
NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation
von: Liu, Youzhi, et al.
Veröffentlicht: (2024)
von: Liu, Youzhi, et al.
Veröffentlicht: (2024)
CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models
von: Song, Wenxuan, et al.
Veröffentlicht: (2026)
von: Song, Wenxuan, et al.
Veröffentlicht: (2026)
SocialNav-SUB: Benchmarking VLMs for Scene Understanding in Social Robot Navigation
von: Munje, Michael J., et al.
Veröffentlicht: (2025)
von: Munje, Michael J., et al.
Veröffentlicht: (2025)
VoroNav: Voronoi-based Zero-shot Object Navigation with Large Language Model
von: Wu, Pengying, et al.
Veröffentlicht: (2024)
von: Wu, Pengying, et al.
Veröffentlicht: (2024)
NavForesee: A Unified Vision-Language World Model for Hierarchical Planning and Dual-Horizon Navigation Prediction
von: Liu, Fei, et al.
Veröffentlicht: (2025)
von: Liu, Fei, et al.
Veröffentlicht: (2025)
SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation
von: Zhang, Jiwen, et al.
Veröffentlicht: (2026)
von: Zhang, Jiwen, et al.
Veröffentlicht: (2026)
AdaVLN: Towards Visual Language Navigation in Continuous Indoor Environments with Moving Humans
von: Loh, Dillon, et al.
Veröffentlicht: (2024)
von: Loh, Dillon, et al.
Veröffentlicht: (2024)
IntentionNav: A Benchmark for Intent-Driven Object Navigation from Implicit Human Instruction
von: Qian, Lin, et al.
Veröffentlicht: (2026)
von: Qian, Lin, et al.
Veröffentlicht: (2026)
Manipulate-Anything: Automating Real-World Robots using Vision-Language Models
von: Duan, Jiafei, et al.
Veröffentlicht: (2024)
von: Duan, Jiafei, et al.
Veröffentlicht: (2024)
NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning
von: Lin, Bingqian, et al.
Veröffentlicht: (2024)
von: Lin, Bingqian, et al.
Veröffentlicht: (2024)
SocialNav-MoE: A Mixture-of-Experts Vision Language Model for Socially Compliant Navigation with Reinforcement Fine-Tuning
von: Kawabata, Tomohito, et al.
Veröffentlicht: (2025)
von: Kawabata, Tomohito, et al.
Veröffentlicht: (2025)
Structured Observation Language for Efficient and Generalizable Vision-Language Navigation
von: Peng, Daojie, et al.
Veröffentlicht: (2026)
von: Peng, Daojie, et al.
Veröffentlicht: (2026)
Enhancing Vision-Language Navigation with Multimodal Event Knowledge from Real-World Indoor Tour Videos
von: Xu, Haoxuan, et al.
Veröffentlicht: (2026)
von: Xu, Haoxuan, et al.
Veröffentlicht: (2026)
Nav-R1: Reasoning and Navigation in Embodied Scenes
von: Liu, Qingxiang, et al.
Veröffentlicht: (2025)
von: Liu, Qingxiang, et al.
Veröffentlicht: (2025)
NavTrust: Benchmarking Trustworthiness for Embodied Navigation
von: Jiang, Huaide, et al.
Veröffentlicht: (2026)
von: Jiang, Huaide, et al.
Veröffentlicht: (2026)
CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model
von: Yu, Zhuoyuan, et al.
Veröffentlicht: (2025)
von: Yu, Zhuoyuan, et al.
Veröffentlicht: (2025)
UrbanNav: Learning Language-Guided Urban Navigation from Web-Scale Human Trajectories
von: Mei, Yanghong, et al.
Veröffentlicht: (2025)
von: Mei, Yanghong, et al.
Veröffentlicht: (2025)
ShadowNav: Autonomous Global Localization for Lunar Navigation in Darkness
von: Atha, Deegan, et al.
Veröffentlicht: (2024)
von: Atha, Deegan, et al.
Veröffentlicht: (2024)
Hyp2Nav: Hyperbolic Planning and Curiosity for Crowd Navigation
von: di Melendugno, Guido Maria D'Amely, et al.
Veröffentlicht: (2024)
von: di Melendugno, Guido Maria D'Amely, et al.
Veröffentlicht: (2024)
FOM-Nav: Frontier-Object Maps for Object Goal Navigation
von: Chabal, Thomas, et al.
Veröffentlicht: (2025)
von: Chabal, Thomas, et al.
Veröffentlicht: (2025)
Towards Realistic UAV Vision-Language Navigation: Platform, Benchmark, and Methodology
von: Wang, Xiangyu, et al.
Veröffentlicht: (2024)
von: Wang, Xiangyu, et al.
Veröffentlicht: (2024)
DreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation
von: Wang, Yunheng, et al.
Veröffentlicht: (2025)
von: Wang, Yunheng, et al.
Veröffentlicht: (2025)
LangNav: Language as a Perceptual Representation for Navigation
von: Pan, Bowen, et al.
Veröffentlicht: (2023)
von: Pan, Bowen, et al.
Veröffentlicht: (2023)
DynaNav: Dynamic Feature and Layer Selection for Efficient Visual Navigation
von: Wang, Jiahui, et al.
Veröffentlicht: (2025)
von: Wang, Jiahui, et al.
Veröffentlicht: (2025)
PIG-Nav: Key Insights for Pretrained Image Goal Navigation Models
von: Wan, Jiansong, et al.
Veröffentlicht: (2025)
von: Wan, Jiansong, et al.
Veröffentlicht: (2025)
OctoNav: Towards Generalist Embodied Navigation
von: Gao, Chen, et al.
Veröffentlicht: (2025)
von: Gao, Chen, et al.
Veröffentlicht: (2025)
Bridging the Indoor-Outdoor Gap: Vision-Centric Instruction-Guided Embodied Navigation for the Last Meters
von: Zhao, Yuxiang, et al.
Veröffentlicht: (2026)
von: Zhao, Yuxiang, et al.
Veröffentlicht: (2026)
RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
von: Yuan, Wentao, et al.
Veröffentlicht: (2024)
von: Yuan, Wentao, et al.
Veröffentlicht: (2024)
IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation
von: Li, Yifan, et al.
Veröffentlicht: (2025)
von: Li, Yifan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
VL-Nav: A Neuro-Symbolic Approach for Reasoning-based Vision-Language Navigation
von: Du, Yi, et al.
Veröffentlicht: (2025) -
NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models
von: Zhou, Gengze, et al.
Veröffentlicht: (2024) -
SignNav: Leveraging Signage for Semantic Visual Navigation in Large-Scale Indoor Environments
von: Sun, Jian, et al.
Veröffentlicht: (2026) -
NavQ: Learning a Q-Model for Foresighted Vision-and-Language Navigation
von: Xu, Peiran, et al.
Veröffentlicht: (2025) -
The One RING: a Robotic Indoor Navigation Generalist
von: Eftekhar, Ainaz, et al.
Veröffentlicht: (2024)