HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Jiaming, Chen, Hao, An, Pengju, Liu, Zhuoyang, Zhang, Renrui, Gu, Chenyang, Li, Xiaoqi, Guo, Ziyu, Chen, Sixiang, Liu, Mengzhen, Hou, Chengkai, Zhao, Mengdi, Zhou, KC alex, Heng, Pheng-Ann, Zhang, Shanghang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915354388725760
author Liu, Jiaming
Chen, Hao
An, Pengju
Liu, Zhuoyang
Zhang, Renrui
Gu, Chenyang
Li, Xiaoqi
Guo, Ziyu
Chen, Sixiang
Liu, Mengzhen
Hou, Chengkai
Zhao, Mengdi
Zhou, KC alex
Heng, Pheng-Ann
Zhang, Shanghang
author_facet Liu, Jiaming
Chen, Hao
An, Pengju
Liu, Zhuoyang
Zhang, Renrui
Gu, Chenyang
Li, Xiaoqi
Guo, Ziyu
Chen, Sixiang
Liu, Mengzhen
Hou, Chengkai
Zhao, Mengdi
Zhou, KC alex
Heng, Pheng-Ann
Zhang, Shanghang
contents A fundamental objective of manipulation policy design is to endow robots to comprehend human instructions, reason about scene cues, and execute generalized actions in dynamic environments. Recent autoregressive vision-language-action (VLA) methods inherit common-sense reasoning capabilities from vision-language models (VLMs) for next action-token prediction. However, these methods quantize actions into discrete bins, which disrupts the continuity required for precise control. In contrast, existing diffusion-based VLA methods incorporate an additional diffusion head to predict continuous actions solely conditioned on feature representations extracted by the VLM, without fully leveraging the VLM's pretrained reasoning capabilities through token-level generation. To address these limitations, we introduce HybridVLA, a unified framework that absorbs the continuous nature of diffusion-based actions and the contextual reasoning of autoregression within a single large language model. To mitigate interference between the two generation paradigms, we propose a collaborative training recipe that seamlessly incorporates diffusion denoising into the next-token prediction process. With this recipe, we find these two action prediction methods not only reinforce each other but also exhibit varying strength across different tasks. Therefore, we design a collaborative action ensemble mechanism that adaptively fuses both predictions, leading to more robust control. HybridVLA outperforms previous state-of-the-art VLA methods by 14\% and 19\% in mean success rate on simulation and real-world tasks, respectively, while demonstrating stable manipulation in unseen configurations.
format Preprint
id arxiv_https___arxiv_org_abs_2503_10631
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
Liu, Jiaming
Chen, Hao
An, Pengju
Liu, Zhuoyang
Zhang, Renrui
Gu, Chenyang
Li, Xiaoqi
Guo, Ziyu
Chen, Sixiang
Liu, Mengzhen
Hou, Chengkai
Zhao, Mengdi
Zhou, KC alex
Heng, Pheng-Ann
Zhang, Shanghang
Computer Vision and Pattern Recognition
Robotics
A fundamental objective of manipulation policy design is to endow robots to comprehend human instructions, reason about scene cues, and execute generalized actions in dynamic environments. Recent autoregressive vision-language-action (VLA) methods inherit common-sense reasoning capabilities from vision-language models (VLMs) for next action-token prediction. However, these methods quantize actions into discrete bins, which disrupts the continuity required for precise control. In contrast, existing diffusion-based VLA methods incorporate an additional diffusion head to predict continuous actions solely conditioned on feature representations extracted by the VLM, without fully leveraging the VLM's pretrained reasoning capabilities through token-level generation. To address these limitations, we introduce HybridVLA, a unified framework that absorbs the continuous nature of diffusion-based actions and the contextual reasoning of autoregression within a single large language model. To mitigate interference between the two generation paradigms, we propose a collaborative training recipe that seamlessly incorporates diffusion denoising into the next-token prediction process. With this recipe, we find these two action prediction methods not only reinforce each other but also exhibit varying strength across different tasks. Therefore, we design a collaborative action ensemble mechanism that adaptively fuses both predictions, leading to more robust control. HybridVLA outperforms previous state-of-the-art VLA methods by 14\% and 19\% in mean success rate on simulation and real-world tasks, respectively, while demonstrating stable manipulation in unseen configurations.
title HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2503.10631