Endowing Embodied Agents with Spatial Reasoning Capabilities for Vision-and-Language Navigation
Fuente:
arXiv
Saved in:
| Main Authors: | Bai, Qianqian, Chen, Zhongpu, Luo, Ling, Du, Huaming, Lei, Yuqian, Jiao, Ziyun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MA-CoNav: A Master-Slave Multi-Agent Framework with Hierarchical Collaboration and Dual-Level Reflection for Long-Horizon Embodied VLN
by: Luo, Ling, et al.
Published: (2026)
by: Luo, Ling, et al.
Published: (2026)
RAGNav: A Retrieval-Augmented Topological Reasoning Framework for Multi-Goal Visual-Language Navigation
by: Luo, Ling, et al.
Published: (2026)
by: Luo, Ling, et al.
Published: (2026)
BioProVLA-Agent: An Affordable, Protocol-Driven, Vision-Enhanced VLA-Enabled Embodied Multi-Agent System with Closed-Loop-Capable Reasoning for Biological Laboratory Manipulation
by: Du, Zhaohui, et al.
Published: (2026)
by: Du, Zhaohui, et al.
Published: (2026)
Exploring Spatial Representation to Enhance LLM Reasoning in Aerial Vision-Language Navigation
by: Gao, Yunpeng, et al.
Published: (2024)
by: Gao, Yunpeng, et al.
Published: (2024)
Provable Ordering and Continuity in Vision-Language Pretraining for Generalizable Embodied Agents
by: Zhang, Zhizhen, et al.
Published: (2025)
by: Zhang, Zhizhen, et al.
Published: (2025)
Vision-Language Navigation with Embodied Intelligence: A Survey
by: Gao, Peng, et al.
Published: (2024)
by: Gao, Peng, et al.
Published: (2024)
Zero-shot Object Navigation with Vision-Language Models Reasoning
by: Wen, Congcong, et al.
Published: (2024)
by: Wen, Congcong, et al.
Published: (2024)
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
by: Chen, Boyuan, et al.
Published: (2024)
by: Chen, Boyuan, et al.
Published: (2024)
Survey of Vision-Language-Action Models for Embodied Manipulation
by: Li, Haoran, et al.
Published: (2025)
by: Li, Haoran, et al.
Published: (2025)
Mem2Ego: Empowering Vision-Language Models with Global-to-Ego Memory for Long-Horizon Embodied Navigation
by: Zhang, Lingfeng, et al.
Published: (2025)
by: Zhang, Lingfeng, et al.
Published: (2025)
SUM-AgriVLN: Spatial Understanding Memory for Agricultural Vision-and-Language Navigation
by: Zhao, Xiaobei, et al.
Published: (2025)
by: Zhao, Xiaobei, et al.
Published: (2025)
Lifelong Embodied Navigation Learning
by: Wang, Xudong, et al.
Published: (2026)
by: Wang, Xudong, et al.
Published: (2026)
A Survey of Robotic Navigation and Manipulation with Physics Simulators in the Era of Embodied AI
by: Wong, Lik Hang Kenny, et al.
Published: (2025)
by: Wong, Lik Hang Kenny, et al.
Published: (2025)
CityNavAgent: Aerial Vision-and-Language Navigation with Hierarchical Semantic Planning and Global Memory
by: Zhang, Weichen, et al.
Published: (2025)
by: Zhang, Weichen, et al.
Published: (2025)
Safety of Embodied Navigation: A Survey
by: Wang, Zixia, et al.
Published: (2025)
by: Wang, Zixia, et al.
Published: (2025)
Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models
by: Lei, Zixing, et al.
Published: (2026)
by: Lei, Zixing, et al.
Published: (2026)
ViLAM: Distilling Vision-Language Reasoning into Attention Maps for Social Robot Navigation
by: Elnoor, Mohamed, et al.
Published: (2025)
by: Elnoor, Mohamed, et al.
Published: (2025)
SpaceMind: A Modular and Self-Evolving Embodied Vision-Language Agent Framework for Autonomous On-orbit Servicing
by: Wu, Aodi, et al.
Published: (2026)
by: Wu, Aodi, et al.
Published: (2026)
VAMOS: A Hierarchical Vision-Language-Action Model for Capability-Modulated and Steerable Navigation
by: Castro, Mateo Guaman, et al.
Published: (2025)
by: Castro, Mateo Guaman, et al.
Published: (2025)
NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models
by: Zhou, Gengze, et al.
Published: (2024)
by: Zhou, Gengze, et al.
Published: (2024)
Vision-Language Navigation with Continual Learning
by: Li, Zhiyuan, et al.
Published: (2024)
by: Li, Zhiyuan, et al.
Published: (2024)
3DGSNav: Enhancing Vision-Language Model Reasoning for Object Navigation via Active 3D Gaussian Splatting
by: Zheng, Wancai, et al.
Published: (2026)
by: Zheng, Wancai, et al.
Published: (2026)
EmboCoach-Bench: Benchmarking AI Agents on Developing Embodied Robots
by: Lei, Zixing, et al.
Published: (2026)
by: Lei, Zixing, et al.
Published: (2026)
Cog-GA: A Large Language Models-based Generative Agent for Vision-Language Navigation in Continuous Environments
by: Li, Zhiyuan, et al.
Published: (2024)
by: Li, Zhiyuan, et al.
Published: (2024)
EmbodiedCity: A Benchmark Platform for Embodied Agent in Real-world City Environment
by: Gao, Chen, et al.
Published: (2024)
by: Gao, Chen, et al.
Published: (2024)
DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models
by: Yin, Cheng, et al.
Published: (2025)
by: Yin, Cheng, et al.
Published: (2025)
ROVER: Recursive Reasoning Over Videos with Vision-Language Models for Embodied Tasks
by: Schroeder, Philip, et al.
Published: (2025)
by: Schroeder, Philip, et al.
Published: (2025)
Mind to Hand: Purposeful Robotic Control via Embodied Reasoning
by: Tang, Peijun, et al.
Published: (2025)
by: Tang, Peijun, et al.
Published: (2025)
Active Test-time Vision-Language Navigation
by: Ko, Heeju, et al.
Published: (2025)
by: Ko, Heeju, et al.
Published: (2025)
VLA-AN: An Efficient and Onboard Vision-Language-Action Framework for Aerial Navigation in Complex Environments
by: Wu, Yuze, et al.
Published: (2025)
by: Wu, Yuze, et al.
Published: (2025)
LatentPilot: Scene-Aware Vision-and-Language Navigation by Dreaming Ahead with Latent Visual Reasoning
by: Hao, Haihong, et al.
Published: (2026)
by: Hao, Haihong, et al.
Published: (2026)
PRISM: A Multi-View Multi-Capability Retail Video Dataset for Embodied Vision-Language Models
by: Rouhi, Amirreza, et al.
Published: (2026)
by: Rouhi, Amirreza, et al.
Published: (2026)
Embodied Spatial Intelligence: from Implicit Scene Modeling to Spatial Reasoning
by: Fang, Jiading
Published: (2025)
by: Fang, Jiading
Published: (2025)
Long-horizon Embodied Planning with Implicit Logical Inference and Hallucination Mitigation
by: Liu, Siyuan, et al.
Published: (2024)
by: Liu, Siyuan, et al.
Published: (2024)
UAV-CodeAgents: Scalable UAV Mission Planning via Multi-Agent ReAct and Vision-Language Reasoning
by: Sautenkov, Oleg, et al.
Published: (2025)
by: Sautenkov, Oleg, et al.
Published: (2025)
PM-Nav: Priori-Map Guided Embodied Navigation in Functional Buildings
by: Gao, Jiang, et al.
Published: (2026)
by: Gao, Jiang, et al.
Published: (2026)
Rethinking the Embodied Gap in Vision-and-Language Navigation: A Holistic Study of Physical and Visual Disparities
by: Wang, Liuyi, et al.
Published: (2025)
by: Wang, Liuyi, et al.
Published: (2025)
Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training
by: Liu, Yi, et al.
Published: (2025)
by: Liu, Yi, et al.
Published: (2025)
TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking
by: Liu, Jiahang, et al.
Published: (2025)
by: Liu, Jiahang, et al.
Published: (2025)
SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation
by: Zhang, Jiwen, et al.
Published: (2026)
by: Zhang, Jiwen, et al.
Published: (2026)
Similar Items
-
MA-CoNav: A Master-Slave Multi-Agent Framework with Hierarchical Collaboration and Dual-Level Reflection for Long-Horizon Embodied VLN
by: Luo, Ling, et al.
Published: (2026) -
RAGNav: A Retrieval-Augmented Topological Reasoning Framework for Multi-Goal Visual-Language Navigation
by: Luo, Ling, et al.
Published: (2026) -
BioProVLA-Agent: An Affordable, Protocol-Driven, Vision-Enhanced VLA-Enabled Embodied Multi-Agent System with Closed-Loop-Capable Reasoning for Biological Laboratory Manipulation
by: Du, Zhaohui, et al.
Published: (2026) -
Exploring Spatial Representation to Enhance LLM Reasoning in Aerial Vision-Language Navigation
by: Gao, Yunpeng, et al.
Published: (2024) -
Provable Ordering and Continuity in Vision-Language Pretraining for Generalizable Embodied Agents
by: Zhang, Zhizhen, et al.
Published: (2025)