Language and Planning in Robotic Navigation: A Multilingual Evaluation of State-of-the-Art Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mansour, Malak, Aly, Ahmed, Tharwat, Bahey, Hashmi, Sarim, An, Dong, Reid, Ian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Latent Action Pretraining Through World Modeling
von: Tharwat, Bahey, et al.
Veröffentlicht: (2025)
von: Tharwat, Bahey, et al.
Veröffentlicht: (2025)
Indexing Multimodal Language Models for Large-scale Image Retrieval
von: Tharwat, Bahey, et al.
Veröffentlicht: (2026)
von: Tharwat, Bahey, et al.
Veröffentlicht: (2026)
ETPNav: Evolving Topological Planning for Vision-Language Navigation in Continuous Environments
von: An, Dong, et al.
Veröffentlicht: (2023)
von: An, Dong, et al.
Veröffentlicht: (2023)
DivScene: Towards Open-Vocabulary Object Navigation with Large Vision Language Models in Diverse Scenes
von: Wang, Zhaowei, et al.
Veröffentlicht: (2024)
von: Wang, Zhaowei, et al.
Veröffentlicht: (2024)
Navigating Beyond Instructions: Vision-and-Language Navigation in Obstructed Environments
von: Hong, Haodong, et al.
Veröffentlicht: (2024)
von: Hong, Haodong, et al.
Veröffentlicht: (2024)
A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges
von: Li, Zongxia, et al.
Veröffentlicht: (2025)
von: Li, Zongxia, et al.
Veröffentlicht: (2025)
EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning
von: Chen, Yi, et al.
Veröffentlicht: (2023)
von: Chen, Yi, et al.
Veröffentlicht: (2023)
End-to-End Navigation with Vision Language Models: Transforming Spatial Reasoning into Question-Answering
von: Goetting, Dylan, et al.
Veröffentlicht: (2024)
von: Goetting, Dylan, et al.
Veröffentlicht: (2024)
Constraint-Aware Zero-Shot Vision-Language Navigation in Continuous Environments
von: Chen, Kehan, et al.
Veröffentlicht: (2024)
von: Chen, Kehan, et al.
Veröffentlicht: (2024)
CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model
von: Yu, Zhuoyuan, et al.
Veröffentlicht: (2025)
von: Yu, Zhuoyuan, et al.
Veröffentlicht: (2025)
Can DeepSeek Reason Like a Surgeon? An Empirical Evaluation for Vision-Language Understanding in Robotic-Assisted Surgery
von: Ma, Boyi, et al.
Veröffentlicht: (2025)
von: Ma, Boyi, et al.
Veröffentlicht: (2025)
RoboUniView: Visual-Language Model with Unified View Representation for Robotic Manipulation
von: Liu, Fanfan, et al.
Veröffentlicht: (2024)
von: Liu, Fanfan, et al.
Veröffentlicht: (2024)
GPT-4V(ision) for Robotics: Multimodal Task Planning from Human Demonstration
von: Wake, Naoki, et al.
Veröffentlicht: (2023)
von: Wake, Naoki, et al.
Veröffentlicht: (2023)
VLN-NF: Feasibility-Aware Vision-and-Language Navigation with False-Premise Instructions
von: Su, Hung-Ting, et al.
Veröffentlicht: (2026)
von: Su, Hung-Ting, et al.
Veröffentlicht: (2026)
Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation
von: Zhang, Pingrui, et al.
Veröffentlicht: (2025)
von: Zhang, Pingrui, et al.
Veröffentlicht: (2025)
Polaris: Open-ended Interactive Robotic Manipulation via Syn2Real Visual Grounding and Large Language Models
von: Wang, Tianyu, et al.
Veröffentlicht: (2024)
von: Wang, Tianyu, et al.
Veröffentlicht: (2024)
What Limits Vision-and-Language Navigation ?
von: Wang, Yunheng, et al.
Veröffentlicht: (2026)
von: Wang, Yunheng, et al.
Veröffentlicht: (2026)
Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation
von: Lu, Jinghui, et al.
Veröffentlicht: (2026)
von: Lu, Jinghui, et al.
Veröffentlicht: (2026)
SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning
von: Schroeder, Philip, et al.
Veröffentlicht: (2026)
von: Schroeder, Philip, et al.
Veröffentlicht: (2026)
Vision-and-Language Navigation Generative Pretrained Transformer
von: Hanlin, Wen
Veröffentlicht: (2024)
von: Hanlin, Wen
Veröffentlicht: (2024)
NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models
von: Zhou, Gengze, et al.
Veröffentlicht: (2024)
von: Zhou, Gengze, et al.
Veröffentlicht: (2024)
From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction
von: Zhao, Zhida, et al.
Veröffentlicht: (2025)
von: Zhao, Zhida, et al.
Veröffentlicht: (2025)
LangNav: Language as a Perceptual Representation for Navigation
von: Pan, Bowen, et al.
Veröffentlicht: (2023)
von: Pan, Bowen, et al.
Veröffentlicht: (2023)
Hierarchical Open-Vocabulary 3D Scene Graphs for Language-Grounded Robot Navigation
von: Werby, Abdelrhman, et al.
Veröffentlicht: (2024)
von: Werby, Abdelrhman, et al.
Veröffentlicht: (2024)
RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation
von: Li, Huiqiong, et al.
Veröffentlicht: (2026)
von: Li, Huiqiong, et al.
Veröffentlicht: (2026)
Unseen from Seen: Rewriting Observation-Instruction Using Foundation Models for Augmenting Vision-Language Navigation
von: Wei, Ziming, et al.
Veröffentlicht: (2025)
von: Wei, Ziming, et al.
Veröffentlicht: (2025)
ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning
von: Yang, Yandan, et al.
Veröffentlicht: (2026)
von: Yang, Yandan, et al.
Veröffentlicht: (2026)
World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning
von: Wang, Siyin, et al.
Veröffentlicht: (2025)
von: Wang, Siyin, et al.
Veröffentlicht: (2025)
Ground-level Viewpoint Vision-and-Language Navigation in Continuous Environments
von: Li, Zerui, et al.
Veröffentlicht: (2025)
von: Li, Zerui, et al.
Veröffentlicht: (2025)
Do Visual Imaginations Improve Vision-and-Language Navigation Agents?
von: Perincherry, Akhil, et al.
Veröffentlicht: (2025)
von: Perincherry, Akhil, et al.
Veröffentlicht: (2025)
Probing Collision Grounding in Vision-Language Models for Safe Human-Robot Collaboration
von: Wang, Jun, et al.
Veröffentlicht: (2026)
von: Wang, Jun, et al.
Veröffentlicht: (2026)
GenSim: Generating Robotic Simulation Tasks via Large Language Models
von: Wang, Lirui, et al.
Veröffentlicht: (2023)
von: Wang, Lirui, et al.
Veröffentlicht: (2023)
From Prompts to Pavement Through Time: Temporal Grounding in Agentic Scene-to-Plan Reasoning
von: Gado, Ahmed Y., et al.
Veröffentlicht: (2026)
von: Gado, Ahmed Y., et al.
Veröffentlicht: (2026)
SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts
von: Zhou, Gengze, et al.
Veröffentlicht: (2024)
von: Zhou, Gengze, et al.
Veröffentlicht: (2024)
NaVILA: Legged Robot Vision-Language-Action Model for Navigation
von: Cheng, An-Chieh, et al.
Veröffentlicht: (2024)
von: Cheng, An-Chieh, et al.
Veröffentlicht: (2024)
InstructNav: Zero-shot System for Generic Instruction Navigation in Unexplored Environment
von: Long, Yuxing, et al.
Veröffentlicht: (2024)
von: Long, Yuxing, et al.
Veröffentlicht: (2024)
RefAV: Towards Planning-Centric Scenario Mining
von: Davidson, Cainan, et al.
Veröffentlicht: (2025)
von: Davidson, Cainan, et al.
Veröffentlicht: (2025)
Real-world Instance-specific Image Goal Navigation: Bridging Domain Gaps via Contrastive Learning
von: Sakaguchi, Taichi, et al.
Veröffentlicht: (2024)
von: Sakaguchi, Taichi, et al.
Veröffentlicht: (2024)
BEVPose: Unveiling Scene Semantics through Pose-Guided Multi-Modal BEV Alignment
von: Hosseinzadeh, Mehdi, et al.
Veröffentlicht: (2024)
von: Hosseinzadeh, Mehdi, et al.
Veröffentlicht: (2024)
A Superalignment Framework in Autonomous Driving with Large Language Models
von: Kong, Xiangrui, et al.
Veröffentlicht: (2024)
von: Kong, Xiangrui, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Latent Action Pretraining Through World Modeling
von: Tharwat, Bahey, et al.
Veröffentlicht: (2025) -
Indexing Multimodal Language Models for Large-scale Image Retrieval
von: Tharwat, Bahey, et al.
Veröffentlicht: (2026) -
ETPNav: Evolving Topological Planning for Vision-Language Navigation in Continuous Environments
von: An, Dong, et al.
Veröffentlicht: (2023) -
DivScene: Towards Open-Vocabulary Object Navigation with Large Vision Language Models in Diverse Scenes
von: Wang, Zhaowei, et al.
Veröffentlicht: (2024) -
Navigating Beyond Instructions: Vision-and-Language Navigation in Obstructed Environments
von: Hong, Haodong, et al.
Veröffentlicht: (2024)