SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Gengze, Hong, Yicong, Wang, Zun, Zhao, Chongyang, Bansal, Mohit, Wu, Qi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models
von: Zhou, Gengze, et al.
Veröffentlicht: (2024)
von: Zhou, Gengze, et al.
Veröffentlicht: (2024)
VLN-MME: Diagnosing MLLMs as Language-guided Visual Navigation agents
von: Zhao, Xunyi, et al.
Veröffentlicht: (2025)
von: Zhao, Xunyi, et al.
Veröffentlicht: (2025)
Ground-level Viewpoint Vision-and-Language Navigation in Continuous Environments
von: Li, Zerui, et al.
Veröffentlicht: (2025)
von: Li, Zerui, et al.
Veröffentlicht: (2025)
Bootstrapping Language-Guided Navigation Learning with Self-Refining Data Flywheel
von: Wang, Zun, et al.
Veröffentlicht: (2024)
von: Wang, Zun, et al.
Veröffentlicht: (2024)
NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation
von: Zhang, Jiazhao, et al.
Veröffentlicht: (2024)
von: Zhang, Jiazhao, et al.
Veröffentlicht: (2024)
Policy-Guided World Model Planning for Language-Conditioned Visual Navigation
von: Chahe, Amirhosein, et al.
Veröffentlicht: (2026)
von: Chahe, Amirhosein, et al.
Veröffentlicht: (2026)
Navigating Beyond Instructions: Vision-and-Language Navigation in Obstructed Environments
von: Hong, Haodong, et al.
Veröffentlicht: (2024)
von: Hong, Haodong, et al.
Veröffentlicht: (2024)
ETPNav: Evolving Topological Planning for Vision-Language Navigation in Continuous Environments
von: An, Dong, et al.
Veröffentlicht: (2023)
von: An, Dong, et al.
Veröffentlicht: (2023)
Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
Rethinking Training Dynamics in Scale-wise Autoregressive Generation
von: Zhou, Gengze, et al.
Veröffentlicht: (2025)
von: Zhou, Gengze, et al.
Veröffentlicht: (2025)
Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation
von: Zhang, Pingrui, et al.
Veröffentlicht: (2025)
von: Zhang, Pingrui, et al.
Veröffentlicht: (2025)
Learning Goal-Oriented Vision-and-Language Navigation with Self-Improving Demonstrations at Scale
von: Li, Songze, et al.
Veröffentlicht: (2025)
von: Li, Songze, et al.
Veröffentlicht: (2025)
Towards Visuospatial Cognition via Hierarchical Fusion of Visual Experts
von: Feng, Qi
Veröffentlicht: (2025)
von: Feng, Qi
Veröffentlicht: (2025)
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
von: Yu, Shoubin, et al.
Veröffentlicht: (2026)
von: Yu, Shoubin, et al.
Veröffentlicht: (2026)
Do Visual Imaginations Improve Vision-and-Language Navigation Agents?
von: Perincherry, Akhil, et al.
Veröffentlicht: (2025)
von: Perincherry, Akhil, et al.
Veröffentlicht: (2025)
DRAGON: A Dialogue-Based Robot for Assistive Navigation with Visual Language Grounding
von: Liu, Shuijing, et al.
Veröffentlicht: (2023)
von: Liu, Shuijing, et al.
Veröffentlicht: (2023)
Affordances-Oriented Planning using Foundation Models for Continuous Vision-Language Navigation
von: Chen, Jiaqi, et al.
Veröffentlicht: (2024)
von: Chen, Jiaqi, et al.
Veröffentlicht: (2024)
Rethinking the Embodied Gap in Vision-and-Language Navigation: A Holistic Study of Physical and Visual Disparities
von: Wang, Liuyi, et al.
Veröffentlicht: (2025)
von: Wang, Liuyi, et al.
Veröffentlicht: (2025)
Selective Exploration and Information Gathering in Search and Rescue Using Hierarchical Learning Guided by Natural Language Input
von: Panagopoulos, Dimitrios, et al.
Veröffentlicht: (2024)
von: Panagopoulos, Dimitrios, et al.
Veröffentlicht: (2024)
Vision-and-Language Navigation Generative Pretrained Transformer
von: Hanlin, Wen
Veröffentlicht: (2024)
von: Hanlin, Wen
Veröffentlicht: (2024)
SmallPlan: Leverage Small Language Models for Sequential Path Planning with Simulation-Powered, LLM-Guided Distillation
von: Pham, Quang P. M., et al.
Veröffentlicht: (2025)
von: Pham, Quang P. M., et al.
Veröffentlicht: (2025)
Skill-Based Mixture-of-Experts: Adaptive Routing for Heterogeneous Reasoning via Inferred Skills
von: Chen, Justin Chih-Yao, et al.
Veröffentlicht: (2025)
von: Chen, Justin Chih-Yao, et al.
Veröffentlicht: (2025)
DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
von: Wang, Zun, et al.
Veröffentlicht: (2024)
von: Wang, Zun, et al.
Veröffentlicht: (2024)
Recursive Visual Imagination and Adaptive Linguistic Grounding for Vision Language Navigation
von: Chen, Bolei, et al.
Veröffentlicht: (2025)
von: Chen, Bolei, et al.
Veröffentlicht: (2025)
OpenFMNav: Towards Open-Set Zero-Shot Object Navigation via Vision-Language Foundation Models
von: Kuang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Kuang, Yuxuan, et al.
Veröffentlicht: (2024)
IPPON: Common Sense Guided Informative Path Planning for Object Goal Navigation
von: Qu, Kaixian, et al.
Veröffentlicht: (2024)
von: Qu, Kaixian, et al.
Veröffentlicht: (2024)
Statler: State-Maintaining Language Models for Embodied Reasoning
von: Yoneda, Takuma, et al.
Veröffentlicht: (2023)
von: Yoneda, Takuma, et al.
Veröffentlicht: (2023)
RACER: Rich Language-Guided Failure Recovery Policies for Imitation Learning
von: Dai, Yinpei, et al.
Veröffentlicht: (2024)
von: Dai, Yinpei, et al.
Veröffentlicht: (2024)
Planning with Sketch-Guided Verification for Physics-Aware Video Generation
von: Huang, Yidong, et al.
Veröffentlicht: (2025)
von: Huang, Yidong, et al.
Veröffentlicht: (2025)
3D-MoE: A Mixture-of-Experts Multi-modal LLM for 3D Vision and Pose Diffusion via Rectified Flow
von: Ma, Yueen, et al.
Veröffentlicht: (2025)
von: Ma, Yueen, et al.
Veröffentlicht: (2025)
PSALM-V: Automating Symbolic Planning in Interactive Visual Environments with Large Language Models
von: Zhu, Wang Bill, et al.
Veröffentlicht: (2025)
von: Zhu, Wang Bill, et al.
Veröffentlicht: (2025)
VLNVerse: A Benchmark for Vision-Language Navigation with Versatile, Embodied, Realistic Simulation and Evaluation
von: Lin, Sihao, et al.
Veröffentlicht: (2025)
von: Lin, Sihao, et al.
Veröffentlicht: (2025)
Leveraging Adaptive Group Negotiation for Heterogeneous Multi-Robot Collaboration with Large Language Models
von: Song, Siqi, et al.
Veröffentlicht: (2025)
von: Song, Siqi, et al.
Veröffentlicht: (2025)
VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation
von: Yu, Shoubin, et al.
Veröffentlicht: (2025)
von: Yu, Shoubin, et al.
Veröffentlicht: (2025)
LGR2: Language Guided Reward Relabeling for Accelerating Hierarchical Reinforcement Learning
von: Singh, Utsav, et al.
Veröffentlicht: (2024)
von: Singh, Utsav, et al.
Veröffentlicht: (2024)
Holodeck: Language Guided Generation of 3D Embodied AI Environments
von: Yang, Yue, et al.
Veröffentlicht: (2023)
von: Yang, Yue, et al.
Veröffentlicht: (2023)
Mind the Error! Detection and Localization of Instruction Errors in Vision-and-Language Navigation
von: Taioli, Francesco, et al.
Veröffentlicht: (2024)
von: Taioli, Francesco, et al.
Veröffentlicht: (2024)
VLN-NF: Feasibility-Aware Vision-and-Language Navigation with False-Premise Instructions
von: Su, Hung-Ting, et al.
Veröffentlicht: (2026)
von: Su, Hung-Ting, et al.
Veröffentlicht: (2026)
What Limits Vision-and-Language Navigation ?
von: Wang, Yunheng, et al.
Veröffentlicht: (2026)
von: Wang, Yunheng, et al.
Veröffentlicht: (2026)
Stable Language Guidance for Vision-Language-Action Models
von: Zhan, Zhihao, et al.
Veröffentlicht: (2026)
von: Zhan, Zhihao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models
von: Zhou, Gengze, et al.
Veröffentlicht: (2024) -
VLN-MME: Diagnosing MLLMs as Language-guided Visual Navigation agents
von: Zhao, Xunyi, et al.
Veröffentlicht: (2025) -
Ground-level Viewpoint Vision-and-Language Navigation in Continuous Environments
von: Li, Zerui, et al.
Veröffentlicht: (2025) -
Bootstrapping Language-Guided Navigation Learning with Self-Refining Data Flywheel
von: Wang, Zun, et al.
Veröffentlicht: (2024) -
NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation
von: Zhang, Jiazhao, et al.
Veröffentlicht: (2024)