SpanVLA: Efficient Action Bridging and Learning from Negative-Recovery Samples for Vision-Language-Action Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Zewei, Yang, Ruining, Xuewei, Qi, Guo, Yiluan, Chen, Sherry X., Feng, Tao, Pistunova, Kateryna, Shen, Yishan, Su, Lili, Ma, Jiaqi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913052033548288
author Zhou, Zewei
Yang, Ruining
Xuewei
Qi
Guo, Yiluan
Chen, Sherry X.
Feng, Tao
Pistunova, Kateryna
Shen, Yishan
Su, Lili
Ma, Jiaqi
author_facet Zhou, Zewei
Yang, Ruining
Xuewei
Qi
Guo, Yiluan
Chen, Sherry X.
Feng, Tao
Pistunova, Kateryna
Shen, Yishan
Su, Lili
Ma, Jiaqi
contents Vision-Language-Action (VLA) models offer a promising autonomous driving paradigm for leveraging world knowledge and reasoning capabilities, especially in long-tail scenarios. However, existing VLA models often struggle with the high latency in action generation using an autoregressive generation framework and exhibit limited robustness. In this paper, we propose SpanVLA, a novel end-to-end autonomous driving framework, integrating an autoregressive reasoning and a flow-matching action expert. First, SpanVLA introduces an efficient bridge to leverage the vision and reasoning guidance of VLM to efficiently plan future trajectories using a flow-matching policy conditioned on historical trajectory initialization, which significantly reduces inference time. Second, to further improve the performance and robustness of the SpanVLA model, we propose a GRPO-based post-training method to enable the VLA model not only to learn from positive driving samples but also to learn how to avoid the typical negative behaviors and learn recovery behaviors. We further introduce mReasoning, a new real-world driving reasoning dataset, focusing on complex, reasoning-demanding scenarios and negative-recovery samples. Extensive experiments on the NAVSIM (v1 and v2) demonstrate the competitive performance of the SpanVLA model. Additionally, the qualitative results across diverse scenarios highlight the planning performance and robustness of our model.
format Preprint
id arxiv_https___arxiv_org_abs_2604_19710
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SpanVLA: Efficient Action Bridging and Learning from Negative-Recovery Samples for Vision-Language-Action Model
Zhou, Zewei
Yang, Ruining
Xuewei
Qi
Guo, Yiluan
Chen, Sherry X.
Feng, Tao
Pistunova, Kateryna
Shen, Yishan
Su, Lili
Ma, Jiaqi
Computer Vision and Pattern Recognition
Vision-Language-Action (VLA) models offer a promising autonomous driving paradigm for leveraging world knowledge and reasoning capabilities, especially in long-tail scenarios. However, existing VLA models often struggle with the high latency in action generation using an autoregressive generation framework and exhibit limited robustness. In this paper, we propose SpanVLA, a novel end-to-end autonomous driving framework, integrating an autoregressive reasoning and a flow-matching action expert. First, SpanVLA introduces an efficient bridge to leverage the vision and reasoning guidance of VLM to efficiently plan future trajectories using a flow-matching policy conditioned on historical trajectory initialization, which significantly reduces inference time. Second, to further improve the performance and robustness of the SpanVLA model, we propose a GRPO-based post-training method to enable the VLA model not only to learn from positive driving samples but also to learn how to avoid the typical negative behaviors and learn recovery behaviors. We further introduce mReasoning, a new real-world driving reasoning dataset, focusing on complex, reasoning-demanding scenarios and negative-recovery samples. Extensive experiments on the NAVSIM (v1 and v2) demonstrate the competitive performance of the SpanVLA model. Additionally, the qualitative results across diverse scenarios highlight the planning performance and robustness of our model.
title SpanVLA: Efficient Action Bridging and Learning from Negative-Recovery Samples for Vision-Language-Action Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.19710