Reasoning-VLA: A Fast and General Vision-Language-Action Reasoning Model for Autonomous Driving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Dapeng, Yuan, Zhenlong, Chen, Zhangquan, Liao, Chih-Ting, Chen, Yinda, Shen, Fei, Zhou, Qingguo, Chua, Tat-Seng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914170153205760
author Zhang, Dapeng
Yuan, Zhenlong
Chen, Zhangquan
Liao, Chih-Ting
Chen, Yinda
Shen, Fei
Zhou, Qingguo
Chua, Tat-Seng
author_facet Zhang, Dapeng
Yuan, Zhenlong
Chen, Zhangquan
Liao, Chih-Ting
Chen, Yinda
Shen, Fei
Zhou, Qingguo
Chua, Tat-Seng
contents Vision-Language-Action (VLA) models have recently shown strong decision-making capabilities in autonomous driving. However, existing VLAs often struggle with achieving efficient inference and generalizing to novel autonomous vehicle configurations and driving scenarios. In this paper, we propose Reasoning-VLA, a general and fast action-generation VLA framework. The proposed model employs a set of learnable action queries, initialized via Gaussian sampling from ground-truth trajectories within the training corpus. These learnable queries interact with reasoning-enhanced vision-language features to generate continuous action trajectories in parallel. To promote robust generalization, we consolidate eight publicly available autonomous driving datasets into a standardized, Chain-of-Thought reasoning-based, and easy-to-use data format for model training. Leveraging both supervised learning and reinforcement learning fine-tuning, extensive empirical evaluations across multiple benchmarks demonstrate that Reasoning-VLA achieves state-of-the-art performance, superior generalization capability, and the excellent inference speed reported to date.
format Preprint
id arxiv_https___arxiv_org_abs_2511_19912
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reasoning-VLA: A Fast and General Vision-Language-Action Reasoning Model for Autonomous Driving
Zhang, Dapeng
Yuan, Zhenlong
Chen, Zhangquan
Liao, Chih-Ting
Chen, Yinda
Shen, Fei
Zhou, Qingguo
Chua, Tat-Seng
Computer Vision and Pattern Recognition
Robotics
Vision-Language-Action (VLA) models have recently shown strong decision-making capabilities in autonomous driving. However, existing VLAs often struggle with achieving efficient inference and generalizing to novel autonomous vehicle configurations and driving scenarios. In this paper, we propose Reasoning-VLA, a general and fast action-generation VLA framework. The proposed model employs a set of learnable action queries, initialized via Gaussian sampling from ground-truth trajectories within the training corpus. These learnable queries interact with reasoning-enhanced vision-language features to generate continuous action trajectories in parallel. To promote robust generalization, we consolidate eight publicly available autonomous driving datasets into a standardized, Chain-of-Thought reasoning-based, and easy-to-use data format for model training. Leveraging both supervised learning and reinforcement learning fine-tuning, extensive empirical evaluations across multiple benchmarks demonstrate that Reasoning-VLA achieves state-of-the-art performance, superior generalization capability, and the excellent inference speed reported to date.
title Reasoning-VLA: A Fast and General Vision-Language-Action Reasoning Model for Autonomous Driving
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2511.19912