Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Wenxuan, Chen, Jiayi, Chen, Shuai, Wang, Jingbo, Ding, Pengxiang, Zhao, Han, Qin, Yikai, Zheng, Xinhu, Wang, Donglin, Wang, Yan, Li, Haoang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913009010475008
author Song, Wenxuan
Chen, Jiayi
Chen, Shuai
Wang, Jingbo
Ding, Pengxiang
Zhao, Han
Qin, Yikai
Zheng, Xinhu
Wang, Donglin
Wang, Yan
Li, Haoang
author_facet Song, Wenxuan
Chen, Jiayi
Chen, Shuai
Wang, Jingbo
Ding, Pengxiang
Zhao, Han
Qin, Yikai
Zheng, Xinhu
Wang, Donglin
Wang, Yan
Li, Haoang
contents This paper proposes a novel approach to address the challenge that pretrained VLA models often fail to effectively improve performance and reduce adaptation costs during standard supervised finetuning (SFT). Some advanced finetuning methods with auxiliary training objectives can improve performance and reduce the number of convergence steps. However, they typically incur significant computational overhead due to the additional losses from auxiliary tasks. To simultaneously achieve the enhanced capabilities of auxiliary training with the simplicity of standard SFT, we decouple the two objectives of auxiliary task training within the parameter space, namely, enhancing general capabilities and fitting task-specific action distributions. To deliver this goal, we only need to train the model to converge on a small-scale task set using two distinct training strategies. The difference between the resulting model parameters can then be interpreted as capability vectors provided by auxiliary tasks. These vectors are then merged with pretrained parameters to form a capability-enhanced meta model. Moreover, when standard SFT is augmented with a lightweight orthogonal regularization loss, the merged model attains performance comparable to auxiliary finetuned baselines with reduced computational overhead. Experimental results demonstrate that this approach is highly effective across diverse robot tasks. Project page: https://chris1220313648.github.io/Fast-dVLA/
format Preprint
id arxiv_https___arxiv_org_abs_2603_25661
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance
Song, Wenxuan
Chen, Jiayi
Chen, Shuai
Wang, Jingbo
Ding, Pengxiang
Zhao, Han
Qin, Yikai
Zheng, Xinhu
Wang, Donglin
Wang, Yan
Li, Haoang
Robotics
Computer Vision and Pattern Recognition
This paper proposes a novel approach to address the challenge that pretrained VLA models often fail to effectively improve performance and reduce adaptation costs during standard supervised finetuning (SFT). Some advanced finetuning methods with auxiliary training objectives can improve performance and reduce the number of convergence steps. However, they typically incur significant computational overhead due to the additional losses from auxiliary tasks. To simultaneously achieve the enhanced capabilities of auxiliary training with the simplicity of standard SFT, we decouple the two objectives of auxiliary task training within the parameter space, namely, enhancing general capabilities and fitting task-specific action distributions. To deliver this goal, we only need to train the model to converge on a small-scale task set using two distinct training strategies. The difference between the resulting model parameters can then be interpreted as capability vectors provided by auxiliary tasks. These vectors are then merged with pretrained parameters to form a capability-enhanced meta model. Moreover, when standard SFT is augmented with a lightweight orthogonal regularization loss, the merged model attains performance comparable to auxiliary finetuned baselines with reduced computational overhead. Experimental results demonstrate that this approach is highly effective across diverse robot tasks. Project page: https://chris1220313648.github.io/Fast-dVLA/
title Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.25661