A Navigation Framework Utilizing Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Duan, Yicheng, tang, Kaiyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
NEBULA: Do We Evaluate Vision-Language-Action Agents Correctly?
von: Peng, Jierui, et al.
Veröffentlicht: (2025)
von: Peng, Jierui, et al.
Veröffentlicht: (2025)
SE-VLN: A Self-Evolving Vision-Language Navigation Framework Based on Multimodal Large Language Models
von: Dong, Xiangyu, et al.
Veröffentlicht: (2025)
von: Dong, Xiangyu, et al.
Veröffentlicht: (2025)
Vision-Language Navigation with Embodied Intelligence: A Survey
von: Gao, Peng, et al.
Veröffentlicht: (2024)
von: Gao, Peng, et al.
Veröffentlicht: (2024)
AgriVLN: Vision-and-Language Navigation for Agricultural Robots
von: Zhao, Xiaobei, et al.
Veröffentlicht: (2025)
von: Zhao, Xiaobei, et al.
Veröffentlicht: (2025)
WorldMAP: Bootstrapping Vision-Language Navigation Trajectory Prediction with Generative World Models
von: Chen, Hongjin, et al.
Veröffentlicht: (2026)
von: Chen, Hongjin, et al.
Veröffentlicht: (2026)
AerialVLA: A Vision-Language-Action Model for UAV Navigation via Minimalist End-to-End Control
von: Xu, Peng, et al.
Veröffentlicht: (2026)
von: Xu, Peng, et al.
Veröffentlicht: (2026)
MapDream: Task-Driven Map Learning for Vision-Language Navigation
von: Lian, Guoxin, et al.
Veröffentlicht: (2026)
von: Lian, Guoxin, et al.
Veröffentlicht: (2026)
What Limits Vision-and-Language Navigation ?
von: Wang, Yunheng, et al.
Veröffentlicht: (2026)
von: Wang, Yunheng, et al.
Veröffentlicht: (2026)
DreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation
von: Wang, Yunheng, et al.
Veröffentlicht: (2025)
von: Wang, Yunheng, et al.
Veröffentlicht: (2025)
RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
von: Yuan, Wentao, et al.
Veröffentlicht: (2024)
von: Yuan, Wentao, et al.
Veröffentlicht: (2024)
Vision-and-Language Navigation Generative Pretrained Transformer
von: Hanlin, Wen
Veröffentlicht: (2024)
von: Hanlin, Wen
Veröffentlicht: (2024)
Dream to Recall: Imagination-Guided Experience Retrieval for Memory-Persistent Vision-and-Language Navigation
von: Xu, Yunzhe, et al.
Veröffentlicht: (2025)
von: Xu, Yunzhe, et al.
Veröffentlicht: (2025)
SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation
von: Zhang, Jiwen, et al.
Veröffentlicht: (2026)
von: Zhang, Jiwen, et al.
Veröffentlicht: (2026)
MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language Navigation
von: Chen, Jiaqi, et al.
Veröffentlicht: (2024)
von: Chen, Jiaqi, et al.
Veröffentlicht: (2024)
Human-Aware Vision-and-Language Navigation: Bridging Simulation to Reality with Dynamic Human Interactions
von: Li, Heng, et al.
Veröffentlicht: (2024)
von: Li, Heng, et al.
Veröffentlicht: (2024)
Language-Conditioned World Modeling for Visual Navigation
von: Dong, Yifei, et al.
Veröffentlicht: (2026)
von: Dong, Yifei, et al.
Veröffentlicht: (2026)
ActiveVLN: Towards Active Exploration via Multi-Turn RL in Vision-and-Language Navigation
von: Zhang, Zekai, et al.
Veröffentlicht: (2025)
von: Zhang, Zekai, et al.
Veröffentlicht: (2025)
LatentPilot: Scene-Aware Vision-and-Language Navigation by Dreaming Ahead with Latent Visual Reasoning
von: Hao, Haihong, et al.
Veröffentlicht: (2026)
von: Hao, Haihong, et al.
Veröffentlicht: (2026)
DINO-CVA: A Multimodal Goal-Conditioned Vision-to-Action Model for Autonomous Catheter Navigation
von: Fekri, Pedram, et al.
Veröffentlicht: (2025)
von: Fekri, Pedram, et al.
Veröffentlicht: (2025)
VLM-RL: A Unified Vision Language Models and Reinforcement Learning Framework for Safe Autonomous Driving
von: Huang, Zilin, et al.
Veröffentlicht: (2024)
von: Huang, Zilin, et al.
Veröffentlicht: (2024)
CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model
von: Yu, Zhuoyuan, et al.
Veröffentlicht: (2025)
von: Yu, Zhuoyuan, et al.
Veröffentlicht: (2025)
NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models
von: Zhou, Gengze, et al.
Veröffentlicht: (2024)
von: Zhou, Gengze, et al.
Veröffentlicht: (2024)
Ground-level Viewpoint Vision-and-Language Navigation in Continuous Environments
von: Li, Zerui, et al.
Veröffentlicht: (2025)
von: Li, Zerui, et al.
Veröffentlicht: (2025)
Do Visual Imaginations Improve Vision-and-Language Navigation Agents?
von: Perincherry, Akhil, et al.
Veröffentlicht: (2025)
von: Perincherry, Akhil, et al.
Veröffentlicht: (2025)
Unseen from Seen: Rewriting Observation-Instruction Using Foundation Models for Augmenting Vision-Language Navigation
von: Wei, Ziming, et al.
Veröffentlicht: (2025)
von: Wei, Ziming, et al.
Veröffentlicht: (2025)
A Survey on Vision-Language-Action Models for Autonomous Driving
von: Jiang, Sicong, et al.
Veröffentlicht: (2025)
von: Jiang, Sicong, et al.
Veröffentlicht: (2025)
Tag Map: A Text-Based Map for Spatial Reasoning and Navigation with Large Language Models
von: Zhang, Mike, et al.
Veröffentlicht: (2024)
von: Zhang, Mike, et al.
Veröffentlicht: (2024)
Exploring the Use of VLMs for Navigation Assistance for People with Blindness and Low Vision
von: Li, Yu, et al.
Veröffentlicht: (2026)
von: Li, Yu, et al.
Veröffentlicht: (2026)
Rethinking the Embodied Gap in Vision-and-Language Navigation: A Holistic Study of Physical and Visual Disparities
von: Wang, Liuyi, et al.
Veröffentlicht: (2025)
von: Wang, Liuyi, et al.
Veröffentlicht: (2025)
UrbanVLA: A Vision-Language-Action Model for Urban Micromobility
von: Li, Anqi, et al.
Veröffentlicht: (2025)
von: Li, Anqi, et al.
Veröffentlicht: (2025)
Physically Grounded Vision-Language Models for Robotic Manipulation
von: Gao, Jensen, et al.
Veröffentlicht: (2023)
von: Gao, Jensen, et al.
Veröffentlicht: (2023)
Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation
von: Zhang, Pingrui, et al.
Veröffentlicht: (2025)
von: Zhang, Pingrui, et al.
Veröffentlicht: (2025)
InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
von: Chen, Xinyi, et al.
Veröffentlicht: (2025)
von: Chen, Xinyi, et al.
Veröffentlicht: (2025)
LIBERO-X: Robustness Litmus for Vision-Language-Action Models
von: Wang, Guodong, et al.
Veröffentlicht: (2026)
von: Wang, Guodong, et al.
Veröffentlicht: (2026)
ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models
von: Tang, Zuojin, et al.
Veröffentlicht: (2026)
von: Tang, Zuojin, et al.
Veröffentlicht: (2026)
A Low-Rank Method for Vision Language Model Hallucination Mitigation in Autonomous Driving
von: Long, Keke, et al.
Veröffentlicht: (2025)
von: Long, Keke, et al.
Veröffentlicht: (2025)
StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing
von: Community, StarVLA
Veröffentlicht: (2026)
von: Community, StarVLA
Veröffentlicht: (2026)
NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning
von: Lin, Bingqian, et al.
Veröffentlicht: (2024)
von: Lin, Bingqian, et al.
Veröffentlicht: (2024)
A Survey on Improving Human Robot Collaboration through Vision-and-Language Navigation
von: Yakolli, Nivedan, et al.
Veröffentlicht: (2025)
von: Yakolli, Nivedan, et al.
Veröffentlicht: (2025)
Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language Navigation
von: Li, Kailing, et al.
Veröffentlicht: (2026)
von: Li, Kailing, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
NEBULA: Do We Evaluate Vision-Language-Action Agents Correctly?
von: Peng, Jierui, et al.
Veröffentlicht: (2025) -
SE-VLN: A Self-Evolving Vision-Language Navigation Framework Based on Multimodal Large Language Models
von: Dong, Xiangyu, et al.
Veröffentlicht: (2025) -
Vision-Language Navigation with Embodied Intelligence: A Survey
von: Gao, Peng, et al.
Veröffentlicht: (2024) -
AgriVLN: Vision-and-Language Navigation for Agricultural Robots
von: Zhao, Xiaobei, et al.
Veröffentlicht: (2025) -
WorldMAP: Bootstrapping Vision-Language Navigation Trajectory Prediction with Generative World Models
von: Chen, Hongjin, et al.
Veröffentlicht: (2026)