MC-GPT: Empowering Vision-and-Language Navigation with Memory Map and Reasoning Chains
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhan, Zhaohuan, Yu, Lisha, Yu, Sijie, Tan, Guang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
EvolveNav: Empowering LLM-Based Vision-Language Navigation via Self-Improving Embodied Reasoning
di: Lin, Bingqian, et al.
Pubblicazione: (2025)
di: Lin, Bingqian, et al.
Pubblicazione: (2025)
NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models
di: Zhou, Gengze, et al.
Pubblicazione: (2024)
di: Zhou, Gengze, et al.
Pubblicazione: (2024)
CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model
di: Yu, Zhuoyuan, et al.
Pubblicazione: (2025)
di: Yu, Zhuoyuan, et al.
Pubblicazione: (2025)
Why Only Text: Empowering Vision-and-Language Navigation with Multi-modal Prompts
di: Hong, Haodong, et al.
Pubblicazione: (2024)
di: Hong, Haodong, et al.
Pubblicazione: (2024)
HopChain: Multi-Hop Data Synthesis for Generalizable Vision-Language Reasoning
di: Wang, Shenzhi, et al.
Pubblicazione: (2026)
di: Wang, Shenzhi, et al.
Pubblicazione: (2026)
Mem4Nav: Boosting Vision-and-Language Navigation in Urban Environments with a Hierarchical Spatial-Cognition Long-Short Memory System
di: He, Lixuan, et al.
Pubblicazione: (2025)
di: He, Lixuan, et al.
Pubblicazione: (2025)
VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning
di: Wang, Qiuchen, et al.
Pubblicazione: (2025)
di: Wang, Qiuchen, et al.
Pubblicazione: (2025)
Navigating the Nuances: A Fine-grained Evaluation of Vision-Language Navigation
di: Wang, Zehao, et al.
Pubblicazione: (2024)
di: Wang, Zehao, et al.
Pubblicazione: (2024)
General Scene Adaptation for Vision-and-Language Navigation
di: Hong, Haodong, et al.
Pubblicazione: (2025)
di: Hong, Haodong, et al.
Pubblicazione: (2025)
Multimodal Chain-of-Thought Reasoning in Language Models
di: Zhang, Zhuosheng, et al.
Pubblicazione: (2023)
di: Zhang, Zhuosheng, et al.
Pubblicazione: (2023)
Seeing Symbols, Missing Cultures: Probing Vision-Language Models' Reasoning on Fire Imagery and Cultural Meaning
di: Yu, Haorui, et al.
Pubblicazione: (2025)
di: Yu, Haorui, et al.
Pubblicazione: (2025)
OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
di: Jia, Mengdi, et al.
Pubblicazione: (2025)
di: Jia, Mengdi, et al.
Pubblicazione: (2025)
NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning
di: Lin, Bingqian, et al.
Pubblicazione: (2024)
di: Lin, Bingqian, et al.
Pubblicazione: (2024)
What Limits Vision-and-Language Navigation ?
di: Wang, Yunheng, et al.
Pubblicazione: (2026)
di: Wang, Yunheng, et al.
Pubblicazione: (2026)
MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language Navigation
di: Chen, Jiaqi, et al.
Pubblicazione: (2024)
di: Chen, Jiaqi, et al.
Pubblicazione: (2024)
ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Model
di: Zhang, Juntian, et al.
Pubblicazione: (2025)
di: Zhang, Juntian, et al.
Pubblicazione: (2025)
DentalGPT: Incentivizing Multimodal Complex Reasoning in Dentistry
di: Cai, Zhenyang, et al.
Pubblicazione: (2025)
di: Cai, Zhenyang, et al.
Pubblicazione: (2025)
Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language Navigation
di: Li, Kailing, et al.
Pubblicazione: (2026)
di: Li, Kailing, et al.
Pubblicazione: (2026)
OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model
di: Chen, Qiguang, et al.
Pubblicazione: (2026)
di: Chen, Qiguang, et al.
Pubblicazione: (2026)
VLN-Video: Utilizing Driving Videos for Outdoor Vision-and-Language Navigation
di: Li, Jialu, et al.
Pubblicazione: (2024)
di: Li, Jialu, et al.
Pubblicazione: (2024)
Correctable Landmark Discovery via Large Models for Vision-Language Navigation
di: Lin, Bingqian, et al.
Pubblicazione: (2024)
di: Lin, Bingqian, et al.
Pubblicazione: (2024)
VLKEB: A Large Vision-Language Model Knowledge Editing Benchmark
di: Huang, Han, et al.
Pubblicazione: (2024)
di: Huang, Han, et al.
Pubblicazione: (2024)
ScanReason: Empowering 3D Visual Grounding with Reasoning Capabilities
di: Zhu, Chenming, et al.
Pubblicazione: (2024)
di: Zhu, Chenming, et al.
Pubblicazione: (2024)
Vision-and-Language Navigation Generative Pretrained Transformer
di: Hanlin, Wen
Pubblicazione: (2024)
di: Hanlin, Wen
Pubblicazione: (2024)
Chain-of-Ground: Improving GUI Grounding via Iterative Reasoning and Reference Feedback
di: Li, Aiden Yiliu, et al.
Pubblicazione: (2025)
di: Li, Aiden Yiliu, et al.
Pubblicazione: (2025)
VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street View
di: Schumann, Raphael, et al.
Pubblicazione: (2023)
di: Schumann, Raphael, et al.
Pubblicazione: (2023)
NoLan: Mitigating Object Hallucinations in Large Vision-Language Models via Dynamic Suppression of Language Priors
di: Ren, Lingfeng, et al.
Pubblicazione: (2026)
di: Ren, Lingfeng, et al.
Pubblicazione: (2026)
On the Cultural Anachronism and Temporal Reasoning in Vision Language Models
di: Ranjan, Mukul, et al.
Pubblicazione: (2026)
di: Ranjan, Mukul, et al.
Pubblicazione: (2026)
Reward Design for Physical Reasoning in Vision-Language Models
di: Lilienthal, Derek, et al.
Pubblicazione: (2026)
di: Lilienthal, Derek, et al.
Pubblicazione: (2026)
Situational Awareness Matters in 3D Vision Language Reasoning
di: Man, Yunze, et al.
Pubblicazione: (2024)
di: Man, Yunze, et al.
Pubblicazione: (2024)
Bootstrapping Language-Guided Navigation Learning with Self-Refining Data Flywheel
di: Wang, Zun, et al.
Pubblicazione: (2024)
di: Wang, Zun, et al.
Pubblicazione: (2024)
DreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation
di: Wang, Yunheng, et al.
Pubblicazione: (2025)
di: Wang, Yunheng, et al.
Pubblicazione: (2025)
EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding
di: Wang, Ziyang, et al.
Pubblicazione: (2026)
di: Wang, Ziyang, et al.
Pubblicazione: (2026)
Semantic Map-based Generation of Navigation Instructions
di: Li, Chengzu, et al.
Pubblicazione: (2024)
di: Li, Chengzu, et al.
Pubblicazione: (2024)
Zero-Shot Visual Reasoning by Vision-Language Models: Benchmarking and Analysis
di: Nagar, Aishik, et al.
Pubblicazione: (2024)
di: Nagar, Aishik, et al.
Pubblicazione: (2024)
Breaking Down and Building Up: Mixture of Skill-Based Vision-and-Language Navigation Agents
di: Ma, Tianyi, et al.
Pubblicazione: (2025)
di: Ma, Tianyi, et al.
Pubblicazione: (2025)
Think Visually, Reason Textually: Vision-Language Synergy in ARC
di: Zhang, Beichen, et al.
Pubblicazione: (2025)
di: Zhang, Beichen, et al.
Pubblicazione: (2025)
ReasonMap: Towards Fine-Grained Visual Reasoning from Transit Maps
di: Feng, Sicheng, et al.
Pubblicazione: (2025)
di: Feng, Sicheng, et al.
Pubblicazione: (2025)
VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning
di: Qi, Yukun, et al.
Pubblicazione: (2025)
di: Qi, Yukun, et al.
Pubblicazione: (2025)
Allegory of the Cave: Measurement-Grounded Vision-Language Learning
di: Xu, Kepeng, et al.
Pubblicazione: (2026)
di: Xu, Kepeng, et al.
Pubblicazione: (2026)
Documenti analoghi
-
EvolveNav: Empowering LLM-Based Vision-Language Navigation via Self-Improving Embodied Reasoning
di: Lin, Bingqian, et al.
Pubblicazione: (2025) -
NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models
di: Zhou, Gengze, et al.
Pubblicazione: (2024) -
CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model
di: Yu, Zhuoyuan, et al.
Pubblicazione: (2025) -
Why Only Text: Empowering Vision-and-Language Navigation with Multi-modal Prompts
di: Hong, Haodong, et al.
Pubblicazione: (2024) -
HopChain: Multi-Hop Data Synthesis for Generalizable Vision-Language Reasoning
di: Wang, Shenzhi, et al.
Pubblicazione: (2026)