MonoDream: Monocular Vision-Language Navigation with Panoramic Dreaming

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Shuo, Wang, Yongcai, Fan, Zhaoxin, Wang, Yucheng, Chen, Maiyue, Wang, Kaihui, Su, Zhizhong, Li, Wanting, Cai, Xudong, Jin, Yeying, Li, Deying
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912733712089088
author Wang, Shuo
Wang, Yongcai
Fan, Zhaoxin
Wang, Yucheng
Chen, Maiyue
Wang, Kaihui
Su, Zhizhong
Li, Wanting
Cai, Xudong
Jin, Yeying
Li, Deying
author_facet Wang, Shuo
Wang, Yongcai
Fan, Zhaoxin
Wang, Yucheng
Chen, Maiyue
Wang, Kaihui
Su, Zhizhong
Li, Wanting
Cai, Xudong
Jin, Yeying
Li, Deying
contents Vision-Language Navigation (VLN) tasks often leverage panoramic RGB and depth inputs to provide rich spatial cues for action planning, but these sensors can be costly or less accessible in real-world deployments. Recent approaches based on Vision-Language Action (VLA) models achieve strong results with monocular input, yet they still lag behind methods using panoramic RGB-D information. We present MonoDream, a lightweight VLA framework that enables monocular agents to learn a Unified Navigation Representation (UNR). This shared feature representation jointly aligns navigation-relevant visual semantics (e.g., global layout, depth, and future cues) and language-grounded action intent, enabling more reliable action prediction. MonoDream further introduces Latent Panoramic Dreaming (LPD) tasks to supervise the UNR, which train the model to predict latent features of panoramic RGB and depth observations at both current and future steps based on only monocular input. Experiments on multiple VLN benchmarks show that MonoDream consistently improves monocular navigation performance and significantly narrows the gap with panoramic-based agents.
format Preprint
id arxiv_https___arxiv_org_abs_2508_02549
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MonoDream: Monocular Vision-Language Navigation with Panoramic Dreaming
Wang, Shuo
Wang, Yongcai
Fan, Zhaoxin
Wang, Yucheng
Chen, Maiyue
Wang, Kaihui
Su, Zhizhong
Li, Wanting
Cai, Xudong
Jin, Yeying
Li, Deying
Computer Vision and Pattern Recognition
Robotics
Vision-Language Navigation (VLN) tasks often leverage panoramic RGB and depth inputs to provide rich spatial cues for action planning, but these sensors can be costly or less accessible in real-world deployments. Recent approaches based on Vision-Language Action (VLA) models achieve strong results with monocular input, yet they still lag behind methods using panoramic RGB-D information. We present MonoDream, a lightweight VLA framework that enables monocular agents to learn a Unified Navigation Representation (UNR). This shared feature representation jointly aligns navigation-relevant visual semantics (e.g., global layout, depth, and future cues) and language-grounded action intent, enabling more reliable action prediction. MonoDream further introduces Latent Panoramic Dreaming (LPD) tasks to supervise the UNR, which train the model to predict latent features of panoramic RGB and depth observations at both current and future steps based on only monocular input. Experiments on multiple VLN benchmarks show that MonoDream consistently improves monocular navigation performance and significantly narrows the gap with panoramic-based agents.
title MonoDream: Monocular Vision-Language Navigation with Panoramic Dreaming
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2508.02549