Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Fan, Yiguo, Ding, Pengxiang, Bai, Shuanghao, Tong, Xinyang, Zhu, Yuyang, Lu, Hongchao, Dai, Fengqi, Zhao, Wei, Liu, Yang, Huang, Siteng, Fan, Zhaoxin, Chen, Badong, Wang, Donglin
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915467564679168
author Fan, Yiguo
Ding, Pengxiang
Bai, Shuanghao
Tong, Xinyang
Zhu, Yuyang
Lu, Hongchao
Dai, Fengqi
Zhao, Wei
Liu, Yang
Huang, Siteng
Fan, Zhaoxin
Chen, Badong
Wang, Donglin
author_facet Fan, Yiguo
Ding, Pengxiang
Bai, Shuanghao
Tong, Xinyang
Zhu, Yuyang
Lu, Hongchao
Dai, Fengqi
Zhao, Wei
Liu, Yang
Huang, Siteng
Fan, Zhaoxin
Chen, Badong
Wang, Donglin
contents Vision-Language-Action (VLA) models have become a cornerstone in robotic policy learning, leveraging large-scale multimodal data for robust and scalable control. However, existing VLA frameworks primarily address short-horizon tasks, and their effectiveness on long-horizon, multi-step robotic manipulation remains limited due to challenges in skill chaining and subtask dependencies. In this work, we introduce Long-VLA, the first end-to-end VLA model specifically designed for long-horizon robotic tasks. Our approach features a novel phase-aware input masking strategy that adaptively segments each subtask into moving and interaction phases, enabling the model to focus on phase-relevant sensory cues and enhancing subtask compatibility. This unified strategy preserves the scalability and data efficiency of VLA training, and our architecture-agnostic module can be seamlessly integrated into existing VLA models. We further propose the L-CALVIN benchmark to systematically evaluate long-horizon manipulation. Extensive experiments on both simulated and real-world tasks demonstrate that Long-VLA significantly outperforms prior state-of-the-art methods, establishing a new baseline for long-horizon robotic control.
format Preprint
id arxiv_https___arxiv_org_abs_2508_19958
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation
Fan, Yiguo
Ding, Pengxiang
Bai, Shuanghao
Tong, Xinyang
Zhu, Yuyang
Lu, Hongchao
Dai, Fengqi
Zhao, Wei
Liu, Yang
Huang, Siteng
Fan, Zhaoxin
Chen, Badong
Wang, Donglin
Robotics
Vision-Language-Action (VLA) models have become a cornerstone in robotic policy learning, leveraging large-scale multimodal data for robust and scalable control. However, existing VLA frameworks primarily address short-horizon tasks, and their effectiveness on long-horizon, multi-step robotic manipulation remains limited due to challenges in skill chaining and subtask dependencies. In this work, we introduce Long-VLA, the first end-to-end VLA model specifically designed for long-horizon robotic tasks. Our approach features a novel phase-aware input masking strategy that adaptively segments each subtask into moving and interaction phases, enabling the model to focus on phase-relevant sensory cues and enhancing subtask compatibility. This unified strategy preserves the scalability and data efficiency of VLA training, and our architecture-agnostic module can be seamlessly integrated into existing VLA models. We further propose the L-CALVIN benchmark to systematically evaluate long-horizon manipulation. Extensive experiments on both simulated and real-world tasks demonstrate that Long-VLA significantly outperforms prior state-of-the-art methods, establishing a new baseline for long-horizon robotic control.
title Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation
topic Robotics
url https://arxiv.org/abs/2508.19958