WMNav: Integrating Vision-Language Models into World Models for Object Goal Navigation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nie, Dujun, Guo, Xianda, Duan, Yiqun, Zhang, Ruijun, Chen, Long |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models
von: Guo, Xianda, et al.
Veröffentlicht: (2024)
von: Guo, Xianda, et al.
Veröffentlicht: (2024)
Goal2Pixel: Grounding Goals to Pixels for Vision-Language Navigation
von: Bao, Muyi, et al.
Veröffentlicht: (2026)
von: Bao, Muyi, et al.
Veröffentlicht: (2026)
WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation
von: Zhao, Baining, et al.
Veröffentlicht: (2026)
von: Zhao, Baining, et al.
Veröffentlicht: (2026)
LARY: A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignment
von: Nie, Dujun, et al.
Veröffentlicht: (2026)
von: Nie, Dujun, et al.
Veröffentlicht: (2026)
FOM-Nav: Frontier-Object Maps for Object Goal Navigation
von: Chabal, Thomas, et al.
Veröffentlicht: (2025)
von: Chabal, Thomas, et al.
Veröffentlicht: (2025)
Cognitive Planning for Object Goal Navigation using Generative AI Models
von: S, Arjun P, et al.
Veröffentlicht: (2024)
von: S, Arjun P, et al.
Veröffentlicht: (2024)
UAV-ON: A Benchmark for Open-World Object Goal Navigation with Aerial Agents
von: Xiao, Jianqiang, et al.
Veröffentlicht: (2025)
von: Xiao, Jianqiang, et al.
Veröffentlicht: (2025)
Language-Based Augmentation to Address Shortcut Learning in Object Goal Navigation
von: Hoftijzer, Dennis, et al.
Veröffentlicht: (2024)
von: Hoftijzer, Dennis, et al.
Veröffentlicht: (2024)
A Navigation Framework Utilizing Vision-Language Models
von: Duan, Yicheng, et al.
Veröffentlicht: (2025)
von: Duan, Yicheng, et al.
Veröffentlicht: (2025)
Distilling LLM Prior to Flow Model for Generalizable Agent's Imagination in Object Goal Navigation
von: Li, Badi, et al.
Veröffentlicht: (2025)
von: Li, Badi, et al.
Veröffentlicht: (2025)
WorldMAP: Bootstrapping Vision-Language Navigation Trajectory Prediction with Generative World Models
von: Chen, Hongjin, et al.
Veröffentlicht: (2026)
von: Chen, Hongjin, et al.
Veröffentlicht: (2026)
DAgger Diffusion Navigation: DAgger Boosted Diffusion Policy for Vision-Language Navigation
von: Shi, Haoxiang, et al.
Veröffentlicht: (2025)
von: Shi, Haoxiang, et al.
Veröffentlicht: (2025)
MOPA: Modular Object Navigation with PointGoal Agents
von: Raychaudhuri, Sonia, et al.
Veröffentlicht: (2023)
von: Raychaudhuri, Sonia, et al.
Veröffentlicht: (2023)
LightStereo: Channel Boost Is All You Need for Efficient 2D Cost Aggregation
von: Guo, Xianda, et al.
Veröffentlicht: (2024)
von: Guo, Xianda, et al.
Veröffentlicht: (2024)
APEX: A Decoupled Memory-based Explorer for Asynchronous Aerial Object Goal Navigation
von: Zhang, Daoxuan, et al.
Veröffentlicht: (2026)
von: Zhang, Daoxuan, et al.
Veröffentlicht: (2026)
DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving
von: jia, Feiyang, et al.
Veröffentlicht: (2026)
von: jia, Feiyang, et al.
Veröffentlicht: (2026)
Sparse Video Generation Propels Real-World Beyond-the-View Vision-Language Navigation
von: Zhang, Hai, et al.
Veröffentlicht: (2026)
von: Zhang, Hai, et al.
Veröffentlicht: (2026)
PIG-Nav: Key Insights for Pretrained Image Goal Navigation Models
von: Wan, Jiansong, et al.
Veröffentlicht: (2025)
von: Wan, Jiansong, et al.
Veröffentlicht: (2025)
Leveraging Large Language Model-based Room-Object Relationships Knowledge for Enhancing Multimodal-Input Object Goal Navigation
von: Sun, Leyuan, et al.
Veröffentlicht: (2024)
von: Sun, Leyuan, et al.
Veröffentlicht: (2024)
VISTAv2: World Imagination for Indoor Vision-and-Language Navigation
von: Huang, Yanjia, et al.
Veröffentlicht: (2025)
von: Huang, Yanjia, et al.
Veröffentlicht: (2025)
TDANet: Target-Directed Attention Network For Object-Goal Visual Navigation With Zero-Shot Ability
von: Lian, Shiwei, et al.
Veröffentlicht: (2024)
von: Lian, Shiwei, et al.
Veröffentlicht: (2024)
Manipulate-Anything: Automating Real-World Robots using Vision-Language Models
von: Duan, Jiafei, et al.
Veröffentlicht: (2024)
von: Duan, Jiafei, et al.
Veröffentlicht: (2024)
POINav: Benchmarking and Enhancing Final-Meters Arrival in Real-World Vision-Language Navigation
von: Gong, Ruiyan, et al.
Veröffentlicht: (2026)
von: Gong, Ruiyan, et al.
Veröffentlicht: (2026)
VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
von: Sun, Jingwen, et al.
Veröffentlicht: (2026)
von: Sun, Jingwen, et al.
Veröffentlicht: (2026)
DOPE: Dual Object Perception-Enhancement Network for Vision-and-Language Navigation
von: Yu, Yinfeng, et al.
Veröffentlicht: (2025)
von: Yu, Yinfeng, et al.
Veröffentlicht: (2025)
Collision-Aware Object-Goal Visual Navigation via Two-Stage Deep Reinforcement Learning
von: Wang, Hongwu, et al.
Veröffentlicht: (2025)
von: Wang, Hongwu, et al.
Veröffentlicht: (2025)
NavForesee: A Unified Vision-Language World Model for Hierarchical Planning and Dual-Horizon Navigation Prediction
von: Liu, Fei, et al.
Veröffentlicht: (2025)
von: Liu, Fei, et al.
Veröffentlicht: (2025)
DivScene: Towards Open-Vocabulary Object Navigation with Large Vision Language Models in Diverse Scenes
von: Wang, Zhaowei, et al.
Veröffentlicht: (2024)
von: Wang, Zhaowei, et al.
Veröffentlicht: (2024)
CapNav: Benchmarking Vision Language Models on Capability-conditioned Indoor Navigation
von: Su, Xia, et al.
Veröffentlicht: (2026)
von: Su, Xia, et al.
Veröffentlicht: (2026)
Zero-shot Object-Centric Instruction Following: Integrating Foundation Models with Traditional Navigation
von: Raychaudhuri, Sonia, et al.
Veröffentlicht: (2024)
von: Raychaudhuri, Sonia, et al.
Veröffentlicht: (2024)
VoroNav: Voronoi-based Zero-shot Object Navigation with Large Language Model
von: Wu, Pengying, et al.
Veröffentlicht: (2024)
von: Wu, Pengying, et al.
Veröffentlicht: (2024)
Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
von: Shao, Rui, et al.
Veröffentlicht: (2025)
von: Shao, Rui, et al.
Veröffentlicht: (2025)
WorldVLM: Combining World Model Forecasting and Vision-Language Reasoning
von: Englmeier, Stefan, et al.
Veröffentlicht: (2026)
von: Englmeier, Stefan, et al.
Veröffentlicht: (2026)
VLD: Visual Language Goal Distance for Reinforcement Learning Navigation
von: Milikic, Lazar, et al.
Veröffentlicht: (2025)
von: Milikic, Lazar, et al.
Veröffentlicht: (2025)
CogNav: Cognitive Process Modeling for Object Goal Navigation with LLMs
von: Cao, Yihan, et al.
Veröffentlicht: (2024)
von: Cao, Yihan, et al.
Veröffentlicht: (2024)
NaVILA: Legged Robot Vision-Language-Action Model for Navigation
von: Cheng, An-Chieh, et al.
Veröffentlicht: (2024)
von: Cheng, An-Chieh, et al.
Veröffentlicht: (2024)
Language-Conditioned World Modeling for Visual Navigation
von: Dong, Yifei, et al.
Veröffentlicht: (2026)
von: Dong, Yifei, et al.
Veröffentlicht: (2026)
BOP-ASK: Object-Interaction Reasoning for Vision-Language Models
von: Bhat, Vineet, et al.
Veröffentlicht: (2025)
von: Bhat, Vineet, et al.
Veröffentlicht: (2025)
DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation
von: Xie, Haozhe, et al.
Veröffentlicht: (2026)
von: Xie, Haozhe, et al.
Veröffentlicht: (2026)
StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling
von: Wei, Meng, et al.
Veröffentlicht: (2025)
von: Wei, Meng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models
von: Guo, Xianda, et al.
Veröffentlicht: (2024) -
Goal2Pixel: Grounding Goals to Pixels for Vision-Language Navigation
von: Bao, Muyi, et al.
Veröffentlicht: (2026) -
WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation
von: Zhao, Baining, et al.
Veröffentlicht: (2026) -
LARY: A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignment
von: Nie, Dujun, et al.
Veröffentlicht: (2026) -
FOM-Nav: Frontier-Object Maps for Object Goal Navigation
von: Chabal, Thomas, et al.
Veröffentlicht: (2025)