NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Gengze, Hong, Yicong, Wang, Zun, Wang, Xin Eric, Wu, Qi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts
von: Zhou, Gengze, et al.
Veröffentlicht: (2024)
von: Zhou, Gengze, et al.
Veröffentlicht: (2024)
Ground-level Viewpoint Vision-and-Language Navigation in Continuous Environments
von: Li, Zerui, et al.
Veröffentlicht: (2025)
von: Li, Zerui, et al.
Veröffentlicht: (2025)
NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation
von: Zhang, Jiazhao, et al.
Veröffentlicht: (2024)
von: Zhang, Jiazhao, et al.
Veröffentlicht: (2024)
ETPNav: Evolving Topological Planning for Vision-Language Navigation in Continuous Environments
von: An, Dong, et al.
Veröffentlicht: (2023)
von: An, Dong, et al.
Veröffentlicht: (2023)
Navigating Beyond Instructions: Vision-and-Language Navigation in Obstructed Environments
von: Hong, Haodong, et al.
Veröffentlicht: (2024)
von: Hong, Haodong, et al.
Veröffentlicht: (2024)
VLN-MME: Diagnosing MLLMs as Language-guided Visual Navigation agents
von: Zhao, Xunyi, et al.
Veröffentlicht: (2025)
von: Zhao, Xunyi, et al.
Veröffentlicht: (2025)
NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning
von: Lin, Bingqian, et al.
Veröffentlicht: (2024)
von: Lin, Bingqian, et al.
Veröffentlicht: (2024)
DreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation
von: Wang, Yunheng, et al.
Veröffentlicht: (2025)
von: Wang, Yunheng, et al.
Veröffentlicht: (2025)
CapNav: Benchmarking Vision Language Models on Capability-conditioned Indoor Navigation
von: Su, Xia, et al.
Veröffentlicht: (2026)
von: Su, Xia, et al.
Veröffentlicht: (2026)
CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model
von: Yu, Zhuoyuan, et al.
Veröffentlicht: (2025)
von: Yu, Zhuoyuan, et al.
Veröffentlicht: (2025)
LangNav: Language as a Perceptual Representation for Navigation
von: Pan, Bowen, et al.
Veröffentlicht: (2023)
von: Pan, Bowen, et al.
Veröffentlicht: (2023)
VL-Nav: A Neuro-Symbolic Approach for Reasoning-based Vision-Language Navigation
von: Du, Yi, et al.
Veröffentlicht: (2025)
von: Du, Yi, et al.
Veröffentlicht: (2025)
Open-Nav: Exploring Zero-Shot Vision-and-Language Navigation in Continuous Environment with Open-Source LLMs
von: Qiao, Yanyuan, et al.
Veröffentlicht: (2024)
von: Qiao, Yanyuan, et al.
Veröffentlicht: (2024)
NavSpace: How Navigation Agents Follow Spatial Intelligence Instructions
von: Yang, Haolin, et al.
Veröffentlicht: (2025)
von: Yang, Haolin, et al.
Veröffentlicht: (2025)
End-to-End Navigation with Vision Language Models: Transforming Spatial Reasoning into Question-Answering
von: Goetting, Dylan, et al.
Veröffentlicht: (2024)
von: Goetting, Dylan, et al.
Veröffentlicht: (2024)
SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation
von: Zhang, Jiwen, et al.
Veröffentlicht: (2026)
von: Zhang, Jiwen, et al.
Veröffentlicht: (2026)
DivScene: Towards Open-Vocabulary Object Navigation with Large Vision Language Models in Diverse Scenes
von: Wang, Zhaowei, et al.
Veröffentlicht: (2024)
von: Wang, Zhaowei, et al.
Veröffentlicht: (2024)
InstructNav: Zero-shot System for Generic Instruction Navigation in Unexplored Environment
von: Long, Yuxing, et al.
Veröffentlicht: (2024)
von: Long, Yuxing, et al.
Veröffentlicht: (2024)
VLNVerse: A Benchmark for Vision-Language Navigation with Versatile, Embodied, Realistic Simulation and Evaluation
von: Lin, Sihao, et al.
Veröffentlicht: (2025)
von: Lin, Sihao, et al.
Veröffentlicht: (2025)
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
von: Chen, Boyuan, et al.
Veröffentlicht: (2024)
von: Chen, Boyuan, et al.
Veröffentlicht: (2024)
What Limits Vision-and-Language Navigation ?
von: Wang, Yunheng, et al.
Veröffentlicht: (2026)
von: Wang, Yunheng, et al.
Veröffentlicht: (2026)
VLN-NF: Feasibility-Aware Vision-and-Language Navigation with False-Premise Instructions
von: Su, Hung-Ting, et al.
Veröffentlicht: (2026)
von: Su, Hung-Ting, et al.
Veröffentlicht: (2026)
VoroNav: Voronoi-based Zero-shot Object Navigation with Large Language Model
von: Wu, Pengying, et al.
Veröffentlicht: (2024)
von: Wu, Pengying, et al.
Veröffentlicht: (2024)
Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation
von: Lu, Jinghui, et al.
Veröffentlicht: (2026)
von: Lu, Jinghui, et al.
Veröffentlicht: (2026)
NavQ: Learning a Q-Model for Foresighted Vision-and-Language Navigation
von: Xu, Peiran, et al.
Veröffentlicht: (2025)
von: Xu, Peiran, et al.
Veröffentlicht: (2025)
Nav-R1: Reasoning and Navigation in Embodied Scenes
von: Liu, Qingxiang, et al.
Veröffentlicht: (2025)
von: Liu, Qingxiang, et al.
Veröffentlicht: (2025)
TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking
von: Liu, Jiahang, et al.
Veröffentlicht: (2025)
von: Liu, Jiahang, et al.
Veröffentlicht: (2025)
Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation
von: Zhang, Pingrui, et al.
Veröffentlicht: (2025)
von: Zhang, Pingrui, et al.
Veröffentlicht: (2025)
VLingNav: Embodied Navigation with Adaptive Reasoning and Visual-Assisted Linguistic Memory
von: Wang, Shaoan, et al.
Veröffentlicht: (2026)
von: Wang, Shaoan, et al.
Veröffentlicht: (2026)
Can DeepSeek Reason Like a Surgeon? An Empirical Evaluation for Vision-Language Understanding in Robotic-Assisted Surgery
von: Ma, Boyi, et al.
Veröffentlicht: (2025)
von: Ma, Boyi, et al.
Veröffentlicht: (2025)
NavForesee: A Unified Vision-Language World Model for Hierarchical Planning and Dual-Horizon Navigation Prediction
von: Liu, Fei, et al.
Veröffentlicht: (2025)
von: Liu, Fei, et al.
Veröffentlicht: (2025)
IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation
von: Li, Yifan, et al.
Veröffentlicht: (2025)
von: Li, Yifan, et al.
Veröffentlicht: (2025)
Rethinking the Embodied Gap in Vision-and-Language Navigation: A Holistic Study of Physical and Visual Disparities
von: Wang, Liuyi, et al.
Veröffentlicht: (2025)
von: Wang, Liuyi, et al.
Veröffentlicht: (2025)
Vision-and-Language Navigation Generative Pretrained Transformer
von: Hanlin, Wen
Veröffentlicht: (2024)
von: Hanlin, Wen
Veröffentlicht: (2024)
MG-Nav: Dual-Scale Visual Navigation via Sparse Spatial Memory
von: Wang, Bo, et al.
Veröffentlicht: (2025)
von: Wang, Bo, et al.
Veröffentlicht: (2025)
Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
HaltNav: Reactive Visual Halting over Lightweight Topological Priors for Robust Vision-Language Navigation
von: Yu, Zihui, et al.
Veröffentlicht: (2026)
von: Yu, Zihui, et al.
Veröffentlicht: (2026)
Three-Step Nav: A Hierarchical Global-Local Planner for Zero-Shot Vision-and-Language Navigation
von: Zheng, Wanrong, et al.
Veröffentlicht: (2026)
von: Zheng, Wanrong, et al.
Veröffentlicht: (2026)
NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation
von: Liu, Youzhi, et al.
Veröffentlicht: (2024)
von: Liu, Youzhi, et al.
Veröffentlicht: (2024)
UrbanNav: Learning Language-Guided Urban Navigation from Web-Scale Human Trajectories
von: Mei, Yanghong, et al.
Veröffentlicht: (2025)
von: Mei, Yanghong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts
von: Zhou, Gengze, et al.
Veröffentlicht: (2024) -
Ground-level Viewpoint Vision-and-Language Navigation in Continuous Environments
von: Li, Zerui, et al.
Veröffentlicht: (2025) -
NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation
von: Zhang, Jiazhao, et al.
Veröffentlicht: (2024) -
ETPNav: Evolving Topological Planning for Vision-Language Navigation in Continuous Environments
von: An, Dong, et al.
Veröffentlicht: (2023) -
Navigating Beyond Instructions: Vision-and-Language Navigation in Obstructed Environments
von: Hong, Haodong, et al.
Veröffentlicht: (2024)