StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Wei, Meng, Wan, Chenyang, Yu, Xiqian, Wang, Tai, Yang, Yuqiang, Mao, Xiaohan, Zhu, Chenming, Cai, Wenzhe, Wang, Hanqing, Chen, Yilun, Liu, Xihui, Pang, Jiangmiao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ground Slow, Move Fast: A Dual-System Foundation Model for Generalizable Vision-and-Language Navigation
by: Wei, Meng, et al.
Published: (2025)
by: Wei, Meng, et al.
Published: (2025)
OVExp: Open Vocabulary Exploration for Object-Oriented Navigation
by: Wei, Meng, et al.
Published: (2024)
by: Wei, Meng, et al.
Published: (2024)
NavDP: Learning Sim-to-Real Navigation Diffusion Policy with Privileged Information Guidance
by: Cai, Wenzhe, et al.
Published: (2025)
by: Cai, Wenzhe, et al.
Published: (2025)
LoGoPlanner: Localization Grounded Navigation Policy with Metric-aware Visual Geometry
by: Peng, Jiaqi, et al.
Published: (2025)
by: Peng, Jiaqi, et al.
Published: (2025)
OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding
by: Lin, Jingli, et al.
Published: (2025)
by: Lin, Jingli, et al.
Published: (2025)
LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
by: Zhu, Chenming, et al.
Published: (2024)
by: Zhu, Chenming, et al.
Published: (2024)
Language-to-Space Programming for Training-Free 3D Visual Grounding
by: Mi, Boyu, et al.
Published: (2025)
by: Mi, Boyu, et al.
Published: (2025)
Towards Latency-Aware 3D Streaming Perception for Autonomous Driving
by: Peng, Jiaqi, et al.
Published: (2025)
by: Peng, Jiaqi, et al.
Published: (2025)
Rethinking the Embodied Gap in Vision-and-Language Navigation: A Holistic Study of Physical and Visual Disparities
by: Wang, Liuyi, et al.
Published: (2025)
by: Wang, Liuyi, et al.
Published: (2025)
SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation
by: Hong, Yining, et al.
Published: (2024)
by: Hong, Yining, et al.
Published: (2024)
FSR-VLN: Fast and Slow Reasoning for Vision-Language Navigation with Hierarchical Multi-modal Scene Graph
by: Zhou, Xiaolin, et al.
Published: (2025)
by: Zhou, Xiaolin, et al.
Published: (2025)
VL-LN Bench: Towards Long-horizon Goal-oriented Navigation with Active Dialogs
by: Huang, Wensi, et al.
Published: (2025)
by: Huang, Wensi, et al.
Published: (2025)
SFMViT: SlowFast Meet ViT in Chaotic World
by: Lin, Jiaying, et al.
Published: (2024)
by: Lin, Jiaying, et al.
Published: (2024)
SlowFast-SCI: Slow-Fast Deep Unfolding Learning for Spectral Compressive Imaging
by: Zeng, Haijin, et al.
Published: (2025)
by: Zeng, Haijin, et al.
Published: (2025)
Slot-VLM: SlowFast Slots for Video-Language Modeling
by: Xu, Jiaqi, et al.
Published: (2024)
by: Xu, Jiaqi, et al.
Published: (2024)
MMScan: A Multi-Modal 3D Scene Dataset with Hierarchical Grounded Language Annotations
by: Lyu, Ruiyuan, et al.
Published: (2024)
by: Lyu, Ruiyuan, et al.
Published: (2024)
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
by: Yang, Shuai, et al.
Published: (2025)
by: Yang, Shuai, et al.
Published: (2025)
Accelerating Diffusion Large Language Models with SlowFast Sampling: The Three Golden Principles
by: Wei, Qingyan, et al.
Published: (2025)
by: Wei, Qingyan, et al.
Published: (2025)
SF-TMN: SlowFast Temporal Modeling Network for Surgical Phase Recognition
by: Zhang, Bokai, et al.
Published: (2023)
by: Zhang, Bokai, et al.
Published: (2023)
Using SlowFast Networks for Near-Miss Incident Analysis in Dashcam Videos
by: Zhang, Yucheng, et al.
Published: (2024)
by: Zhang, Yucheng, et al.
Published: (2024)
Streaming, Fast and Slow: Cognitive Load-Aware Streaming for Efficient LLM Serving
by: Xiao, Chang, et al.
Published: (2025)
by: Xiao, Chang, et al.
Published: (2025)
WhisperNetV2: SlowFast Siamese Network For Lip-Based Biometrics
by: Zakeri, Abdollah, et al.
Published: (2024)
by: Zakeri, Abdollah, et al.
Published: (2024)
PhysMamba: Efficient Remote Physiological Measurement with SlowFast Temporal Difference Mamba
by: Luo, Chaoqi, et al.
Published: (2024)
by: Luo, Chaoqi, et al.
Published: (2024)
Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM
by: Huang, Haifeng, et al.
Published: (2026)
by: Huang, Haifeng, et al.
Published: (2026)
CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
AgentVLN: Towards Agentic Vision-and-Language Navigation
by: Xin, Zihao, et al.
Published: (2026)
by: Xin, Zihao, et al.
Published: (2026)
PointLLM: Empowering Large Language Models to Understand Point Clouds
by: Xu, Runsen, et al.
Published: (2023)
by: Xu, Runsen, et al.
Published: (2023)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
by: Xu, Runsen, et al.
Published: (2024)
by: Xu, Runsen, et al.
Published: (2024)
ScanReason: Empowering 3D Visual Grounding with Reasoning Capabilities
by: Zhu, Chenming, et al.
Published: (2024)
by: Zhu, Chenming, et al.
Published: (2024)
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
by: Xu, Mingze, et al.
Published: (2024)
by: Xu, Mingze, et al.
Published: (2024)
LiveVLN: Breaking the Stop-and-Go Loop in Vision-Language Navigation
by: Wang, Xiangchen, et al.
Published: (2026)
by: Wang, Xiangchen, et al.
Published: (2026)
Efficient-VLN: A Training-Efficient Vision-Language Navigation Model
by: Zheng, Duo, et al.
Published: (2025)
by: Zheng, Duo, et al.
Published: (2025)
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
by: Huang, Haifeng, et al.
Published: (2025)
by: Huang, Haifeng, et al.
Published: (2025)
AgriVLN: Vision-and-Language Navigation for Agricultural Robots
by: Zhao, Xiaobei, et al.
Published: (2025)
by: Zhao, Xiaobei, et al.
Published: (2025)
FlexVLN: Flexible Adaptation for Diverse Vision-and-Language Navigation Tasks
by: Zhang, Siqi, et al.
Published: (2025)
by: Zhang, Siqi, et al.
Published: (2025)
GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation
by: Gao, Ning, et al.
Published: (2025)
by: Gao, Ning, et al.
Published: (2025)
DecoVLN: Decoupling Observation, Reasoning, and Correction for Vision-and-Language Navigation
by: Xin, Zihao, et al.
Published: (2026)
by: Xin, Zihao, et al.
Published: (2026)
G$^2$VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
by: Hu, Wenbo, et al.
Published: (2025)
by: Hu, Wenbo, et al.
Published: (2025)
Modulating State Space Model with SlowFast Framework for Compute-Efficient Ultra Low-Latency Speech Enhancement
by: Cheng, Longbiao, et al.
Published: (2024)
by: Cheng, Longbiao, et al.
Published: (2024)
InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts
by: Zhong, Weipeng, et al.
Published: (2025)
by: Zhong, Weipeng, et al.
Published: (2025)
Similar Items
-
Ground Slow, Move Fast: A Dual-System Foundation Model for Generalizable Vision-and-Language Navigation
by: Wei, Meng, et al.
Published: (2025) -
OVExp: Open Vocabulary Exploration for Object-Oriented Navigation
by: Wei, Meng, et al.
Published: (2024) -
NavDP: Learning Sim-to-Real Navigation Diffusion Policy with Privileged Information Guidance
by: Cai, Wenzhe, et al.
Published: (2025) -
LoGoPlanner: Localization Grounded Navigation Policy with Metric-aware Visual Geometry
by: Peng, Jiaqi, et al.
Published: (2025) -
OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding
by: Lin, Jingli, et al.
Published: (2025)