StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wei, Meng, Wan, Chenyang, Yu, Xiqian, Wang, Tai, Yang, Yuqiang, Mao, Xiaohan, Zhu, Chenming, Cai, Wenzhe, Wang, Hanqing, Chen, Yilun, Liu, Xihui, Pang, Jiangmiao
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908438227845120
author Wei, Meng
Wan, Chenyang
Yu, Xiqian
Wang, Tai
Yang, Yuqiang
Mao, Xiaohan
Zhu, Chenming
Cai, Wenzhe
Wang, Hanqing
Chen, Yilun
Liu, Xihui
Pang, Jiangmiao
author_facet Wei, Meng
Wan, Chenyang
Yu, Xiqian
Wang, Tai
Yang, Yuqiang
Mao, Xiaohan
Zhu, Chenming
Cai, Wenzhe
Wang, Hanqing
Chen, Yilun
Liu, Xihui
Pang, Jiangmiao
contents Vision-and-Language Navigation (VLN) in real-world settings requires agents to process continuous visual streams and generate actions with low latency grounded in language instructions. While Video-based Large Language Models (Video-LLMs) have driven recent progress, current VLN methods based on Video-LLM often face trade-offs among fine-grained visual understanding, long-term context modeling and computational efficiency. We introduce StreamVLN, a streaming VLN framework that employs a hybrid slow-fast context modeling strategy to support multi-modal reasoning over interleaved vision, language and action inputs. The fast-streaming dialogue context facilitates responsive action generation through a sliding-window of active dialogues, while the slow-updating memory context compresses historical visual states using a 3D-aware token pruning strategy. With this slow-fast design, StreamVLN achieves coherent multi-turn dialogue through efficient KV cache reuse, supporting long video streams with bounded context size and inference cost. Experiments on VLN-CE benchmarks demonstrate state-of-the-art performance with stable low latency, ensuring robustness and efficiency in real-world deployment. The project page is: \href{https://streamvln.github.io/}{https://streamvln.github.io/}.
format Preprint
id arxiv_https___arxiv_org_abs_2507_05240
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling
Wei, Meng
Wan, Chenyang
Yu, Xiqian
Wang, Tai
Yang, Yuqiang
Mao, Xiaohan
Zhu, Chenming
Cai, Wenzhe
Wang, Hanqing
Chen, Yilun
Liu, Xihui
Pang, Jiangmiao
Robotics
Computer Vision and Pattern Recognition
Vision-and-Language Navigation (VLN) in real-world settings requires agents to process continuous visual streams and generate actions with low latency grounded in language instructions. While Video-based Large Language Models (Video-LLMs) have driven recent progress, current VLN methods based on Video-LLM often face trade-offs among fine-grained visual understanding, long-term context modeling and computational efficiency. We introduce StreamVLN, a streaming VLN framework that employs a hybrid slow-fast context modeling strategy to support multi-modal reasoning over interleaved vision, language and action inputs. The fast-streaming dialogue context facilitates responsive action generation through a sliding-window of active dialogues, while the slow-updating memory context compresses historical visual states using a 3D-aware token pruning strategy. With this slow-fast design, StreamVLN achieves coherent multi-turn dialogue through efficient KV cache reuse, supporting long video streams with bounded context size and inference cost. Experiments on VLN-CE benchmarks demonstrate state-of-the-art performance with stable low latency, ensuring robustness and efficiency in real-world deployment. The project page is: \href{https://streamvln.github.io/}{https://streamvln.github.io/}.
title StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.05240