Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN
Fuente:
arXiv
Salvato in:
| Autori principali: | Xia, Ziyi, Xiong, Chaoran, Wei, Litao, Hu, Xinhao, Pei, Ling |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SFCo-Nav: Efficient Zero-Shot Visual Language Navigation via Collaboration of Slow LLM and Fast Attributed Graph Alignment
di: Xiong, Chaoran, et al.
Pubblicazione: (2026)
di: Xiong, Chaoran, et al.
Pubblicazione: (2026)
VLN-Game: Vision-Language Equilibrium Search for Zero-Shot Semantic Navigation
di: Yu, Bangguo, et al.
Pubblicazione: (2024)
di: Yu, Bangguo, et al.
Pubblicazione: (2024)
Sensing, Social, and Motion Intelligence in Embodied Navigation: A Comprehensive Survey
di: Xiong, Chaoran, et al.
Pubblicazione: (2025)
di: Xiong, Chaoran, et al.
Pubblicazione: (2025)
HiMemVLN: Enhancing Reliability of Open-Source Zero-Shot Vision-and-Language Navigation with Hierarchical Memory System
di: Lyu, Kailin, et al.
Pubblicazione: (2026)
di: Lyu, Kailin, et al.
Pubblicazione: (2026)
FSR-VLN: Fast and Slow Reasoning for Vision-Language Navigation with Hierarchical Multi-modal Scene Graph
di: Zhou, Xiaolin, et al.
Pubblicazione: (2025)
di: Zhou, Xiaolin, et al.
Pubblicazione: (2025)
Think, Remember, Navigate: Zero-Shot Object-Goal Navigation with VLM-Powered Reasoning
di: Habibpour, Mobin, et al.
Pubblicazione: (2025)
di: Habibpour, Mobin, et al.
Pubblicazione: (2025)
THE-SEAN: A Heart Rate Variation-Inspired Temporally High-Order Event-Based Visual Odometry with Self-Supervised Spiking Event Accumulation Networks
di: Xiong, Chaoran, et al.
Pubblicazione: (2025)
di: Xiong, Chaoran, et al.
Pubblicazione: (2025)
DV-VLN: Dual Verification for Reliable LLM-Based Vision-and-Language Navigation
di: Li, Zijun, et al.
Pubblicazione: (2026)
di: Li, Zijun, et al.
Pubblicazione: (2026)
SUM-AgriVLN: Spatial Understanding Memory for Agricultural Vision-and-Language Navigation
di: Zhao, Xiaobei, et al.
Pubblicazione: (2025)
di: Zhao, Xiaobei, et al.
Pubblicazione: (2025)
VLN-Zero: Rapid Exploration and Cache-Enabled Neurosymbolic Vision-Language Planning for Zero-Shot Transfer in Robot Navigation
di: Bhatt, Neel P., et al.
Pubblicazione: (2025)
di: Bhatt, Neel P., et al.
Pubblicazione: (2025)
AgentVLN: Towards Agentic Vision-and-Language Navigation
di: Xin, Zihao, et al.
Pubblicazione: (2026)
di: Xin, Zihao, et al.
Pubblicazione: (2026)
VLN-Cache: Enabling Token Caching for VLN Models with Visual/Semantic Dynamics Awareness
di: Zheng, Zihao, et al.
Pubblicazione: (2026)
di: Zheng, Zihao, et al.
Pubblicazione: (2026)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
di: Xu, Runsen, et al.
Pubblicazione: (2024)
di: Xu, Runsen, et al.
Pubblicazione: (2024)
OpenVLN: Open-world Aerial Vision-Language Navigation
di: Lin, Peican, et al.
Pubblicazione: (2025)
di: Lin, Peican, et al.
Pubblicazione: (2025)
DyNaVLM: Zero-Shot Vision-Language Navigation System with Dynamic Viewpoints and Self-Refining Graph Memory
di: Ji, Zihe, et al.
Pubblicazione: (2025)
di: Ji, Zihe, et al.
Pubblicazione: (2025)
OmniVLN: Omnidirectional 3D Perception and Token-Efficient LLM Reasoning for Visual-Language Navigation across Air and Ground Platforms
di: Liu, Zhongyuang, et al.
Pubblicazione: (2026)
di: Liu, Zhongyuang, et al.
Pubblicazione: (2026)
T-araVLN: Translator for Agricultural Robotic Agents on Vision-and-Language Navigation
di: Zhao, Xiaobei, et al.
Pubblicazione: (2025)
di: Zhao, Xiaobei, et al.
Pubblicazione: (2025)
LiveVLN: Breaking the Stop-and-Go Loop in Vision-Language Navigation
di: Wang, Xiangchen, et al.
Pubblicazione: (2026)
di: Wang, Xiangchen, et al.
Pubblicazione: (2026)
DecoVLN: Decoupling Observation, Reasoning, and Correction for Vision-and-Language Navigation
di: Xin, Zihao, et al.
Pubblicazione: (2026)
di: Xin, Zihao, et al.
Pubblicazione: (2026)
JanusVLN: Decoupling Semantics and Spatiality with Dual Implicit Memory for Vision-Language Navigation
di: Zeng, Shuang, et al.
Pubblicazione: (2025)
di: Zeng, Shuang, et al.
Pubblicazione: (2025)
AdaVLN: Towards Visual Language Navigation in Continuous Indoor Environments with Moving Humans
di: Loh, Dillon, et al.
Pubblicazione: (2024)
di: Loh, Dillon, et al.
Pubblicazione: (2024)
USS-Nav: Unified Spatio-Semantic Scene Graph for Lightweight UAV Zero-Shot Object Navigation
di: Gai, Weiqi, et al.
Pubblicazione: (2026)
di: Gai, Weiqi, et al.
Pubblicazione: (2026)
MSGNav: Unleashing the Power of Multi-modal 3D Scene Graph for Zero-Shot Embodied Navigation
di: Huang, Xun, et al.
Pubblicazione: (2025)
di: Huang, Xun, et al.
Pubblicazione: (2025)
NavDreamer: Video Models as Zero-Shot 3D Navigators
di: Huang, Xijie, et al.
Pubblicazione: (2026)
di: Huang, Xijie, et al.
Pubblicazione: (2026)
SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation
di: Zhang, Jiwen, et al.
Pubblicazione: (2026)
di: Zhang, Jiwen, et al.
Pubblicazione: (2026)
FlexVLN: Flexible Adaptation for Diverse Vision-and-Language Navigation Tasks
di: Zhang, Siqi, et al.
Pubblicazione: (2025)
di: Zhang, Siqi, et al.
Pubblicazione: (2025)
MDE-AgriVLN: Agricultural Vision-and-Language Navigation with Monocular Depth Estimation
di: Zhao, Xiaobei, et al.
Pubblicazione: (2025)
di: Zhao, Xiaobei, et al.
Pubblicazione: (2025)
SEDualVLN: A Spatially-Enhanced Dual-System for Vision-Language Navigation
di: Huang, Jingzhi, et al.
Pubblicazione: (2026)
di: Huang, Jingzhi, et al.
Pubblicazione: (2026)
VLM-Auto: VLM-based Autonomous Driving Assistant with Human-like Behavior and Understanding for Complex Road Scenes
di: Guo, Ziang, et al.
Pubblicazione: (2024)
di: Guo, Ziang, et al.
Pubblicazione: (2024)
BeliefMapNav: 3D Voxel-Based Belief Map for Zero-Shot Object Navigation
di: Zhou, Zibo, et al.
Pubblicazione: (2025)
di: Zhou, Zibo, et al.
Pubblicazione: (2025)
Exploring the Reliability of Foundation Model-Based Frontier Selection in Zero-Shot Object Goal Navigation
di: Yuan, Shuaihang, et al.
Pubblicazione: (2024)
di: Yuan, Shuaihang, et al.
Pubblicazione: (2024)
IMOST: Incremental Memory Mechanism with Online Self-Supervision for Continual Traversability Learning
di: Ma, Kehui, et al.
Pubblicazione: (2024)
di: Ma, Kehui, et al.
Pubblicazione: (2024)
A2I-Calib: An Anti-noise Active Multi-IMU Spatial-temporal Calibration Framework for Legged Robots
di: Xiong, Chaoran, et al.
Pubblicazione: (2025)
di: Xiong, Chaoran, et al.
Pubblicazione: (2025)
M-SEVIQ: A Multi-band Stereo Event Visual-Inertial Quadruped-based Dataset for Perception under Rapid Motion and Challenging Illumination
di: Cao, Jingcheng, et al.
Pubblicazione: (2026)
di: Cao, Jingcheng, et al.
Pubblicazione: (2026)
DyGeoVLN: Infusing Dynamic Geometry Foundation Model into Vision-Language Navigation
di: Liu, Xiangchen, et al.
Pubblicazione: (2026)
di: Liu, Xiangchen, et al.
Pubblicazione: (2026)
GC-VLN: Instruction as Graph Constraints for Training-free Vision-and-Language Navigation
di: Yin, Hang, et al.
Pubblicazione: (2025)
di: Yin, Hang, et al.
Pubblicazione: (2025)
AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation
di: Guo, Wenxuan, et al.
Pubblicazione: (2026)
di: Guo, Wenxuan, et al.
Pubblicazione: (2026)
Augmenting Tactile Simulators with Real-like and Zero-Shot Capabilities
di: Azulay, Osher, et al.
Pubblicazione: (2023)
di: Azulay, Osher, et al.
Pubblicazione: (2023)
VLA-3D: A Dataset for 3D Semantic Scene Understanding and Navigation
di: Zhang, Haochen, et al.
Pubblicazione: (2024)
di: Zhang, Haochen, et al.
Pubblicazione: (2024)
Multi-Floor Zero-Shot Object Navigation Policy
di: Zhang, Lingfeng, et al.
Pubblicazione: (2024)
di: Zhang, Lingfeng, et al.
Pubblicazione: (2024)
Documenti analoghi
-
SFCo-Nav: Efficient Zero-Shot Visual Language Navigation via Collaboration of Slow LLM and Fast Attributed Graph Alignment
di: Xiong, Chaoran, et al.
Pubblicazione: (2026) -
VLN-Game: Vision-Language Equilibrium Search for Zero-Shot Semantic Navigation
di: Yu, Bangguo, et al.
Pubblicazione: (2024) -
Sensing, Social, and Motion Intelligence in Embodied Navigation: A Comprehensive Survey
di: Xiong, Chaoran, et al.
Pubblicazione: (2025) -
HiMemVLN: Enhancing Reliability of Open-Source Zero-Shot Vision-and-Language Navigation with Hierarchical Memory System
di: Lyu, Kailin, et al.
Pubblicazione: (2026) -
FSR-VLN: Fast and Slow Reasoning for Vision-Language Navigation with Hierarchical Multi-modal Scene Graph
di: Zhou, Xiaolin, et al.
Pubblicazione: (2025)