Vision-and-Language Navigation Generative Pretrained Transformer
Fuente:
arXiv
Salvato in:
| Autore principale: | Hanlin, Wen |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
What Limits Vision-and-Language Navigation ?
di: Wang, Yunheng, et al.
Pubblicazione: (2026)
di: Wang, Yunheng, et al.
Pubblicazione: (2026)
Ground-level Viewpoint Vision-and-Language Navigation in Continuous Environments
di: Li, Zerui, et al.
Pubblicazione: (2025)
di: Li, Zerui, et al.
Pubblicazione: (2025)
Do Visual Imaginations Improve Vision-and-Language Navigation Agents?
di: Perincherry, Akhil, et al.
Pubblicazione: (2025)
di: Perincherry, Akhil, et al.
Pubblicazione: (2025)
NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models
di: Zhou, Gengze, et al.
Pubblicazione: (2024)
di: Zhou, Gengze, et al.
Pubblicazione: (2024)
CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model
di: Yu, Zhuoyuan, et al.
Pubblicazione: (2025)
di: Yu, Zhuoyuan, et al.
Pubblicazione: (2025)
Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation
di: Zhang, Pingrui, et al.
Pubblicazione: (2025)
di: Zhang, Pingrui, et al.
Pubblicazione: (2025)
NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning
di: Lin, Bingqian, et al.
Pubblicazione: (2024)
di: Lin, Bingqian, et al.
Pubblicazione: (2024)
Rethinking the Embodied Gap in Vision-and-Language Navigation: A Holistic Study of Physical and Visual Disparities
di: Wang, Liuyi, et al.
Pubblicazione: (2025)
di: Wang, Liuyi, et al.
Pubblicazione: (2025)
DreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation
di: Wang, Yunheng, et al.
Pubblicazione: (2025)
di: Wang, Yunheng, et al.
Pubblicazione: (2025)
Unseen from Seen: Rewriting Observation-Instruction Using Foundation Models for Augmenting Vision-Language Navigation
di: Wei, Ziming, et al.
Pubblicazione: (2025)
di: Wei, Ziming, et al.
Pubblicazione: (2025)
Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language Navigation
di: Li, Kailing, et al.
Pubblicazione: (2026)
di: Li, Kailing, et al.
Pubblicazione: (2026)
LangNav: Language as a Perceptual Representation for Navigation
di: Pan, Bowen, et al.
Pubblicazione: (2023)
di: Pan, Bowen, et al.
Pubblicazione: (2023)
TRAVEL: Training-Free Retrieval and Alignment for Vision-and-Language Navigation
di: Rajabi, Navid, et al.
Pubblicazione: (2025)
di: Rajabi, Navid, et al.
Pubblicazione: (2025)
Meanings and Measurements: Multi-Agent Probabilistic Grounding for Vision-Language Navigation
di: Padhan, Swagat, et al.
Pubblicazione: (2026)
di: Padhan, Swagat, et al.
Pubblicazione: (2026)
3D-VLA: A 3D Vision-Language-Action Generative World Model
di: Zhen, Haoyu, et al.
Pubblicazione: (2024)
di: Zhen, Haoyu, et al.
Pubblicazione: (2024)
InstructNav: Zero-shot System for Generic Instruction Navigation in Unexplored Environment
di: Long, Yuxing, et al.
Pubblicazione: (2024)
di: Long, Yuxing, et al.
Pubblicazione: (2024)
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
di: Zhou, Zhongyi, et al.
Pubblicazione: (2025)
di: Zhou, Zhongyi, et al.
Pubblicazione: (2025)
Probing Collision Grounding in Vision-Language Models for Safe Human-Robot Collaboration
di: Wang, Jun, et al.
Pubblicazione: (2026)
di: Wang, Jun, et al.
Pubblicazione: (2026)
ROVER: Recursive Reasoning Over Videos with Vision-Language Models for Embodied Tasks
di: Schroeder, Philip, et al.
Pubblicazione: (2025)
di: Schroeder, Philip, et al.
Pubblicazione: (2025)
Explore and Explain: Self-supervised Navigation and Recounting
di: Bigazzi, Roberto, et al.
Pubblicazione: (2020)
di: Bigazzi, Roberto, et al.
Pubblicazione: (2020)
FLAME: Learning to Navigate with Multimodal LLM in Urban Environments
di: Xu, Yunzhe, et al.
Pubblicazione: (2024)
di: Xu, Yunzhe, et al.
Pubblicazione: (2024)
WorldMAP: Bootstrapping Vision-Language Navigation Trajectory Prediction with Generative World Models
di: Chen, Hongjin, et al.
Pubblicazione: (2026)
di: Chen, Hongjin, et al.
Pubblicazione: (2026)
RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
di: Song, Chan Hee, et al.
Pubblicazione: (2024)
di: Song, Chan Hee, et al.
Pubblicazione: (2024)
SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts
di: Zhou, Gengze, et al.
Pubblicazione: (2024)
di: Zhou, Gengze, et al.
Pubblicazione: (2024)
MAPS: Preserving Vision-Language Representations via Module-Wise Proximity Scheduling for Better Vision-Language-Action Generalization
di: Huang, Chengyue, et al.
Pubblicazione: (2025)
di: Huang, Chengyue, et al.
Pubblicazione: (2025)
Vision-Language Navigation with Embodied Intelligence: A Survey
di: Gao, Peng, et al.
Pubblicazione: (2024)
di: Gao, Peng, et al.
Pubblicazione: (2024)
A Navigation Framework Utilizing Vision-Language Models
di: Duan, Yicheng, et al.
Pubblicazione: (2025)
di: Duan, Yicheng, et al.
Pubblicazione: (2025)
AgriVLN: Vision-and-Language Navigation for Agricultural Robots
di: Zhao, Xiaobei, et al.
Pubblicazione: (2025)
di: Zhao, Xiaobei, et al.
Pubblicazione: (2025)
NavSpace: How Navigation Agents Follow Spatial Intelligence Instructions
di: Yang, Haolin, et al.
Pubblicazione: (2025)
di: Yang, Haolin, et al.
Pubblicazione: (2025)
Following Route Instructions using Large Vision-Language Models: A Comparison between Low-level and Panoramic Action Spaces
di: Kåsene, Vebjørn Haug, et al.
Pubblicazione: (2025)
di: Kåsene, Vebjørn Haug, et al.
Pubblicazione: (2025)
General Scene Adaptation for Vision-and-Language Navigation
di: Hong, Haodong, et al.
Pubblicazione: (2025)
di: Hong, Haodong, et al.
Pubblicazione: (2025)
Holodeck: Language Guided Generation of 3D Embodied AI Environments
di: Yang, Yue, et al.
Pubblicazione: (2023)
di: Yang, Yue, et al.
Pubblicazione: (2023)
HERO: Hierarchical Traversable 3D Scene Graphs for Embodied Navigation Among Movable Obstacles
di: Wang, Yunheng, et al.
Pubblicazione: (2025)
di: Wang, Yunheng, et al.
Pubblicazione: (2025)
Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
di: Grover, Shresth, et al.
Pubblicazione: (2025)
di: Grover, Shresth, et al.
Pubblicazione: (2025)
MapDream: Task-Driven Map Learning for Vision-Language Navigation
di: Lian, Guoxin, et al.
Pubblicazione: (2026)
di: Lian, Guoxin, et al.
Pubblicazione: (2026)
End-to-End Navigation with Vision Language Models: Transforming Spatial Reasoning into Question-Answering
di: Goetting, Dylan, et al.
Pubblicazione: (2024)
di: Goetting, Dylan, et al.
Pubblicazione: (2024)
Language and Planning in Robotic Navigation: A Multilingual Evaluation of State-of-the-Art Models
di: Mansour, Malak, et al.
Pubblicazione: (2025)
di: Mansour, Malak, et al.
Pubblicazione: (2025)
Navigating Beyond Instructions: Vision-and-Language Navigation in Obstructed Environments
di: Hong, Haodong, et al.
Pubblicazione: (2024)
di: Hong, Haodong, et al.
Pubblicazione: (2024)
MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language Navigation
di: Chen, Jiaqi, et al.
Pubblicazione: (2024)
di: Chen, Jiaqi, et al.
Pubblicazione: (2024)
Human-Aware Vision-and-Language Navigation: Bridging Simulation to Reality with Dynamic Human Interactions
di: Li, Heng, et al.
Pubblicazione: (2024)
di: Li, Heng, et al.
Pubblicazione: (2024)
Documenti analoghi
-
What Limits Vision-and-Language Navigation ?
di: Wang, Yunheng, et al.
Pubblicazione: (2026) -
Ground-level Viewpoint Vision-and-Language Navigation in Continuous Environments
di: Li, Zerui, et al.
Pubblicazione: (2025) -
Do Visual Imaginations Improve Vision-and-Language Navigation Agents?
di: Perincherry, Akhil, et al.
Pubblicazione: (2025) -
NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models
di: Zhou, Gengze, et al.
Pubblicazione: (2024) -
CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model
di: Yu, Zhuoyuan, et al.
Pubblicazione: (2025)