Zero-shot Object Navigation with Vision-Language Models Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Wen, Congcong, Huang, Yisiyuan, Huang, Hao, Huang, Yanjia, Yuan, Shuaihang, Hao, Yu, Lin, Hui, Liu, Yu-Shen, Fang, Yi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reliable Semantic Understanding for Real World Zero-shot Object Goal Navigation
by: Unlu, Halil Utku, et al.
Published: (2024)
by: Unlu, Halil Utku, et al.
Published: (2024)
GAMap: Zero-Shot Object Goal Navigation with Multi-Scale Geometric-Affordance Guidance
by: Yuan, Shuaihang, et al.
Published: (2024)
by: Yuan, Shuaihang, et al.
Published: (2024)
Exploring the Reliability of Foundation Model-Based Frontier Selection in Zero-Shot Object Goal Navigation
by: Yuan, Shuaihang, et al.
Published: (2024)
by: Yuan, Shuaihang, et al.
Published: (2024)
How Secure Are Large Language Models (LLMs) for Navigation in Urban Environments?
by: Wen, Congcong, et al.
Published: (2024)
by: Wen, Congcong, et al.
Published: (2024)
Humanoid Agent via Embodied Chain-of-Action Reasoning with Multimodal Foundation Models for Zero-Shot Loco-Manipulation
by: Wen, Congcong, et al.
Published: (2025)
by: Wen, Congcong, et al.
Published: (2025)
One-shot Adaptation of Humanoid Whole-body Motion with Walking Priors
by: Huang, Hao, et al.
Published: (2025)
by: Huang, Hao, et al.
Published: (2025)
Socially-Aware Robot Navigation Enhanced by Bidirectional Natural Language Conversations Using Large Language Models
by: Wen, Congcong, et al.
Published: (2024)
by: Wen, Congcong, et al.
Published: (2024)
A Chain-of-Thought Subspace Meta-Learning for Few-shot Image Captioning with Large Vision and Language Models
by: Huang, Hao, et al.
Published: (2025)
by: Huang, Hao, et al.
Published: (2025)
Wavelet Policy: Lifting Scheme for Policy Learning in Long-Horizon Tasks
by: Huang, Hao, et al.
Published: (2025)
by: Huang, Hao, et al.
Published: (2025)
3DGSNav: Enhancing Vision-Language Model Reasoning for Object Navigation via Active 3D Gaussian Splatting
by: Zheng, Wancai, et al.
Published: (2026)
by: Zheng, Wancai, et al.
Published: (2026)
VISTA: Generative Visual Imagination for Vision-and-Language Navigation
by: Huang, Yanjia, et al.
Published: (2025)
by: Huang, Yanjia, et al.
Published: (2025)
MapBERT: Bitwise Masked Modeling for Real-Time Semantic Mapping Generation
by: Deng, Yijie, et al.
Published: (2025)
by: Deng, Yijie, et al.
Published: (2025)
Schrödinger's Navigator: Imagining an Ensemble of Futures for Zero-Shot Object Navigation
by: He, Yu, et al.
Published: (2025)
by: He, Yu, et al.
Published: (2025)
History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation
by: Habibpour, Mobin, et al.
Published: (2025)
by: Habibpour, Mobin, et al.
Published: (2025)
H2-COMPACT: Human-Humanoid Co-Manipulation via Adaptive Contact Trajectory Policies
by: Bethala, Geeta Chandra Raju, et al.
Published: (2025)
by: Bethala, Geeta Chandra Raju, et al.
Published: (2025)
VISTAv2: World Imagination for Indoor Vision-and-Language Navigation
by: Huang, Yanjia, et al.
Published: (2025)
by: Huang, Yanjia, et al.
Published: (2025)
Conversational Orientation Reasoning: Egocentric-to-Allocentric Navigation with Multimodal Chain-of-Thought
by: Huang, Yu Ti
Published: (2025)
by: Huang, Yu Ti
Published: (2025)
AnyImageNav: Any-View Geometry for Precise Last-Meter Image-Goal Navigation
by: Deng, Yijie, et al.
Published: (2026)
by: Deng, Yijie, et al.
Published: (2026)
Think, Remember, Navigate: Zero-Shot Object-Goal Navigation with VLM-Powered Reasoning
by: Habibpour, Mobin, et al.
Published: (2025)
by: Habibpour, Mobin, et al.
Published: (2025)
TopV-Nav: Unlocking the Top-View Spatial Reasoning Potential of MLLM for Zero-shot Object Navigation
by: Zhong, Linqing, et al.
Published: (2024)
by: Zhong, Linqing, et al.
Published: (2024)
VLAD-Grasp: Zero-shot Grasp Detection via Vision-Language Models
by: Kulshrestha, Manav, et al.
Published: (2025)
by: Kulshrestha, Manav, et al.
Published: (2025)
Co-NavGPT: Multi-Robot Cooperative Visual Semantic Navigation Using Vision Language Models
by: Yu, Bangguo, et al.
Published: (2023)
by: Yu, Bangguo, et al.
Published: (2023)
VoroNav: Voronoi-based Zero-shot Object Navigation with Large Language Model
by: Wu, Pengying, et al.
Published: (2024)
by: Wu, Pengying, et al.
Published: (2024)
One Map to Find Them All: Real-time Open-Vocabulary Mapping for Zero-shot Multi-Object Navigation
by: Busch, Finn Lukas, et al.
Published: (2024)
by: Busch, Finn Lukas, et al.
Published: (2024)
AudioScene: Integrating Object-Event Audio into 3D Scenes
by: Yuan, Shuaihang, et al.
Published: (2025)
by: Yuan, Shuaihang, et al.
Published: (2025)
LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transfer
by: Zha, Lihan, et al.
Published: (2026)
by: Zha, Lihan, et al.
Published: (2026)
SemNav: A Model-Based Planner for Zero-Shot Object Goal Navigation Using Vision-Foundation Models
by: Debnath, Arnab, et al.
Published: (2025)
by: Debnath, Arnab, et al.
Published: (2025)
Hierarchical Scoring with 3D Gaussian Splatting for Instance Image-Goal Navigation
by: Deng, Yijie, et al.
Published: (2025)
by: Deng, Yijie, et al.
Published: (2025)
InstructNav: Zero-shot System for Generic Instruction Navigation in Unexplored Environment
by: Long, Yuxing, et al.
Published: (2024)
by: Long, Yuxing, et al.
Published: (2024)
Nav-EE: Navigation-Guided Early Exiting for Efficient Vision-Language Models in Autonomous Driving
by: Hu, Haibo, et al.
Published: (2025)
by: Hu, Haibo, et al.
Published: (2025)
IndoorUAV: Benchmarking Vision-Language UAV Navigation in Continuous Indoor Environments
by: Liu, Xu, et al.
Published: (2025)
by: Liu, Xu, et al.
Published: (2025)
Integrating Retrospective Framework in Multi-Robot Collaboration
by: Liang, Jiazhao, et al.
Published: (2025)
by: Liang, Jiazhao, et al.
Published: (2025)
Thinking in Text and Images: Interleaved Vision--Language Reasoning Traces for Long-Horizon Robot Manipulation
by: Liu, Jinkun, et al.
Published: (2026)
by: Liu, Jinkun, et al.
Published: (2026)
Mem2Ego: Empowering Vision-Language Models with Global-to-Ego Memory for Long-Horizon Embodied Navigation
by: Zhang, Lingfeng, et al.
Published: (2025)
by: Zhang, Lingfeng, et al.
Published: (2025)
ControlVLA: Few-shot Object-centric Adaptation for Pre-trained Vision-Language-Action Models
by: Li, Puhao, et al.
Published: (2025)
by: Li, Puhao, et al.
Published: (2025)
Endowing Embodied Agents with Spatial Reasoning Capabilities for Vision-and-Language Navigation
by: Bai, Qianqian, et al.
Published: (2025)
by: Bai, Qianqian, et al.
Published: (2025)
Unified Control Framework for Real-Time Interception and Obstacle Avoidance of Fast-Moving Objects with Diffusion Variational Autoencoder
by: Dastider, Apan, et al.
Published: (2022)
by: Dastider, Apan, et al.
Published: (2022)
GarmentPile++: Affordance-Driven Cluttered Garments Retrieval with Vision-Language Reasoning
by: Li, Mingleyang, et al.
Published: (2026)
by: Li, Mingleyang, et al.
Published: (2026)
Grounded Vision-Language Navigation for UAVs with Open-Vocabulary Goal Understanding
by: Zhang, Yuhang, et al.
Published: (2025)
by: Zhang, Yuhang, et al.
Published: (2025)
HyPerNav: Hybrid Perception for Object-Oriented Navigation in Unknown Environment
by: Yin, Zecheng, et al.
Published: (2025)
by: Yin, Zecheng, et al.
Published: (2025)
Similar Items
-
Reliable Semantic Understanding for Real World Zero-shot Object Goal Navigation
by: Unlu, Halil Utku, et al.
Published: (2024) -
GAMap: Zero-Shot Object Goal Navigation with Multi-Scale Geometric-Affordance Guidance
by: Yuan, Shuaihang, et al.
Published: (2024) -
Exploring the Reliability of Foundation Model-Based Frontier Selection in Zero-Shot Object Goal Navigation
by: Yuan, Shuaihang, et al.
Published: (2024) -
How Secure Are Large Language Models (LLMs) for Navigation in Urban Environments?
by: Wen, Congcong, et al.
Published: (2024) -
Humanoid Agent via Embodied Chain-of-Action Reasoning with Multimodal Foundation Models for Zero-Shot Loco-Manipulation
by: Wen, Congcong, et al.
Published: (2025)