Reasoning-VLA: A Fast and General Vision-Language-Action Reasoning Model for Autonomous Driving
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914170153205760 |
|---|---|
| author | Zhang, Dapeng Yuan, Zhenlong Chen, Zhangquan Liao, Chih-Ting Chen, Yinda Shen, Fei Zhou, Qingguo Chua, Tat-Seng |
| author_facet | Zhang, Dapeng Yuan, Zhenlong Chen, Zhangquan Liao, Chih-Ting Chen, Yinda Shen, Fei Zhou, Qingguo Chua, Tat-Seng |
| contents | Vision-Language-Action (VLA) models have recently shown strong decision-making capabilities in autonomous driving. However, existing VLAs often struggle with achieving efficient inference and generalizing to novel autonomous vehicle configurations and driving scenarios. In this paper, we propose Reasoning-VLA, a general and fast action-generation VLA framework. The proposed model employs a set of learnable action queries, initialized via Gaussian sampling from ground-truth trajectories within the training corpus. These learnable queries interact with reasoning-enhanced vision-language features to generate continuous action trajectories in parallel. To promote robust generalization, we consolidate eight publicly available autonomous driving datasets into a standardized, Chain-of-Thought reasoning-based, and easy-to-use data format for model training. Leveraging both supervised learning and reinforcement learning fine-tuning, extensive empirical evaluations across multiple benchmarks demonstrate that Reasoning-VLA achieves state-of-the-art performance, superior generalization capability, and the excellent inference speed reported to date. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_19912 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Reasoning-VLA: A Fast and General Vision-Language-Action Reasoning Model for Autonomous Driving Zhang, Dapeng Yuan, Zhenlong Chen, Zhangquan Liao, Chih-Ting Chen, Yinda Shen, Fei Zhou, Qingguo Chua, Tat-Seng Computer Vision and Pattern Recognition Robotics Vision-Language-Action (VLA) models have recently shown strong decision-making capabilities in autonomous driving. However, existing VLAs often struggle with achieving efficient inference and generalizing to novel autonomous vehicle configurations and driving scenarios. In this paper, we propose Reasoning-VLA, a general and fast action-generation VLA framework. The proposed model employs a set of learnable action queries, initialized via Gaussian sampling from ground-truth trajectories within the training corpus. These learnable queries interact with reasoning-enhanced vision-language features to generate continuous action trajectories in parallel. To promote robust generalization, we consolidate eight publicly available autonomous driving datasets into a standardized, Chain-of-Thought reasoning-based, and easy-to-use data format for model training. Leveraging both supervised learning and reinforcement learning fine-tuning, extensive empirical evaluations across multiple benchmarks demonstrate that Reasoning-VLA achieves state-of-the-art performance, superior generalization capability, and the excellent inference speed reported to date. |
| title | Reasoning-VLA: A Fast and General Vision-Language-Action Reasoning Model for Autonomous Driving |
| topic | Computer Vision and Pattern Recognition Robotics |
| url | https://arxiv.org/abs/2511.19912 |