PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Wenxuan, Chen, Jiayi, Ding, Pengxiang, Zhao, Han, Zhao, Wei, Zhong, Zhide, Ge, Zongyuan, Li, Zhijun, Wang, Donglin, Ma, Jun, Wang, Lujia, Li, Haoang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908851378323456
author Song, Wenxuan
Chen, Jiayi
Ding, Pengxiang
Zhao, Han
Zhao, Wei
Zhong, Zhide
Ge, Zongyuan
Li, Zhijun
Wang, Donglin
Ma, Jun
Wang, Lujia
Li, Haoang
author_facet Song, Wenxuan
Chen, Jiayi
Ding, Pengxiang
Zhao, Han
Zhao, Wei
Zhong, Zhide
Ge, Zongyuan
Li, Zhijun
Wang, Donglin
Ma, Jun
Wang, Lujia
Li, Haoang
contents Vision-Language-Action (VLA) models demonstrate remarkable potential for generalizable robotic manipulation. The performance of VLA models can be improved by integrating with action chunking, a critical technique for effective control. However, action chunking linearly scales up action dimensions in VLA models with increased chunking sizes. This reduces the inference efficiency. To tackle this problem, we propose PD-VLA, the first parallel decoding framework for VLA models integrated with action chunking. Our framework reformulates autoregressive decoding as a nonlinear system solved by parallel fixed-point iterations. This approach preserves model performance with mathematical guarantees while significantly improving decoding speed. In addition, it enables training-free acceleration without architectural changes, as well as seamless synergy with existing acceleration techniques. Extensive simulations validate that our PD-VLA maintains competitive success rates while achieving 2.52 times execution frequency on manipulators (with 7 degrees of freedom) compared with the fundamental VLA model. Furthermore, we experimentally identify the most effective settings for acceleration. Finally, real-world experiments validate its high applicability across different tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2503_02310
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding
Song, Wenxuan
Chen, Jiayi
Ding, Pengxiang
Zhao, Han
Zhao, Wei
Zhong, Zhide
Ge, Zongyuan
Li, Zhijun
Wang, Donglin
Ma, Jun
Wang, Lujia
Li, Haoang
Robotics
Computer Vision and Pattern Recognition
Vision-Language-Action (VLA) models demonstrate remarkable potential for generalizable robotic manipulation. The performance of VLA models can be improved by integrating with action chunking, a critical technique for effective control. However, action chunking linearly scales up action dimensions in VLA models with increased chunking sizes. This reduces the inference efficiency. To tackle this problem, we propose PD-VLA, the first parallel decoding framework for VLA models integrated with action chunking. Our framework reformulates autoregressive decoding as a nonlinear system solved by parallel fixed-point iterations. This approach preserves model performance with mathematical guarantees while significantly improving decoding speed. In addition, it enables training-free acceleration without architectural changes, as well as seamless synergy with existing acceleration techniques. Extensive simulations validate that our PD-VLA maintains competitive success rates while achieving 2.52 times execution frequency on manipulators (with 7 degrees of freedom) compared with the fundamental VLA model. Furthermore, we experimentally identify the most effective settings for acceleration. Finally, real-world experiments validate its high applicability across different tasks.
title PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.02310