FantasyVLN: Unified Multimodal Chain-of-Thought Reasoning for Vision-Language Navigation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zuo, Jing, Mu, Lingzhou, Jiang, Fan, Ma, Chengcheng, Xu, Mu, Qi, Yonggang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908783151677440
author Zuo, Jing
Mu, Lingzhou
Jiang, Fan
Ma, Chengcheng
Xu, Mu
Qi, Yonggang
author_facet Zuo, Jing
Mu, Lingzhou
Jiang, Fan
Ma, Chengcheng
Xu, Mu
Qi, Yonggang
contents Achieving human-level performance in Vision-and-Language Navigation (VLN) requires an embodied agent to jointly understand multimodal instructions and visual-spatial context while reasoning over long action sequences. Recent works, such as NavCoT and NavGPT-2, demonstrate the potential of Chain-of-Thought (CoT) reasoning for improving interpretability and long-horizon planning. Moreover, multimodal extensions like OctoNav-R1 and CoT-VLA further validate CoT as a promising pathway toward human-like navigation reasoning. However, existing approaches face critical drawbacks: purely textual CoTs lack spatial grounding and easily overfit to sparse annotated reasoning steps, while multimodal CoTs incur severe token inflation by generating imagined visual observations, making real-time navigation impractical. In this work, we propose FantasyVLN, a unified implicit reasoning framework that preserves the benefits of CoT reasoning without explicit token overhead. Specifically, imagined visual tokens are encoded into a compact latent space using a pretrained Visual AutoRegressor (VAR) during CoT reasoning training, and the model jointly learns from textual, visual, and multimodal CoT modes under a unified multi-CoT strategy. At inference, our model performs direct instruction-to-action mapping while still enjoying reasoning-aware representations. Extensive experiments on LH-VLN show that our approach achieves reasoning-aware yet real-time navigation, improving success rates and efficiency while reducing inference latency by an order of magnitude compared to explicit CoT methods.
format Preprint
id arxiv_https___arxiv_org_abs_2601_13976
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FantasyVLN: Unified Multimodal Chain-of-Thought Reasoning for Vision-Language Navigation
Zuo, Jing
Mu, Lingzhou
Jiang, Fan
Ma, Chengcheng
Xu, Mu
Qi, Yonggang
Computer Vision and Pattern Recognition
Robotics
Achieving human-level performance in Vision-and-Language Navigation (VLN) requires an embodied agent to jointly understand multimodal instructions and visual-spatial context while reasoning over long action sequences. Recent works, such as NavCoT and NavGPT-2, demonstrate the potential of Chain-of-Thought (CoT) reasoning for improving interpretability and long-horizon planning. Moreover, multimodal extensions like OctoNav-R1 and CoT-VLA further validate CoT as a promising pathway toward human-like navigation reasoning. However, existing approaches face critical drawbacks: purely textual CoTs lack spatial grounding and easily overfit to sparse annotated reasoning steps, while multimodal CoTs incur severe token inflation by generating imagined visual observations, making real-time navigation impractical. In this work, we propose FantasyVLN, a unified implicit reasoning framework that preserves the benefits of CoT reasoning without explicit token overhead. Specifically, imagined visual tokens are encoded into a compact latent space using a pretrained Visual AutoRegressor (VAR) during CoT reasoning training, and the model jointly learns from textual, visual, and multimodal CoT modes under a unified multi-CoT strategy. At inference, our model performs direct instruction-to-action mapping while still enjoying reasoning-aware representations. Extensive experiments on LH-VLN show that our approach achieves reasoning-aware yet real-time navigation, improving success rates and efficiency while reducing inference latency by an order of magnitude compared to explicit CoT methods.
title FantasyVLN: Unified Multimodal Chain-of-Thought Reasoning for Vision-Language Navigation
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2601.13976