dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wen, Junjie, Zhu, Minjie, Liu, Jiaming, Liu, Zhiyuan, Yang, Yicun, Zhang, Linfeng, Zhang, Shanghang, Zhu, Yichen, Xu, Yi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914067104399360
author Wen, Junjie
Zhu, Minjie
Liu, Jiaming
Liu, Zhiyuan
Yang, Yicun
Zhang, Linfeng
Zhang, Shanghang
Zhu, Yichen
Xu, Yi
author_facet Wen, Junjie
Zhu, Minjie
Liu, Jiaming
Liu, Zhiyuan
Yang, Yicun
Zhang, Linfeng
Zhang, Shanghang
Zhu, Yichen
Xu, Yi
contents Vision-Language-Action (VLA) models are emerging as a next-generation paradigm for robotics. We introduce dVLA, a diffusion-based VLA that leverages a multimodal chain-of-thought to unify visual perception, language reasoning, and robotic control in a single system. dVLA jointly optimizes perception, language understanding, and action under a single diffusion objective, enabling stronger cross-modal reasoning and better generalization to novel instructions and objects. For practical deployment, we mitigate inference latency by incorporating two acceleration strategies, a prefix attention mask and KV caching, yielding up to around times speedup at test-time inference. We evaluate dVLA in both simulation and the real world: on the LIBERO benchmark, it achieves state-of-the-art performance with a 96.4% average success rate, consistently surpassing both discrete and continuous action policies; on a real Franka robot, it succeeds across a diverse task suite, including a challenging bin-picking task that requires multi-step planning, demonstrating robust real-world performance. Together, these results underscore the promise of unified diffusion frameworks for practical, high-performance VLA robotics.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25681
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
Wen, Junjie
Zhu, Minjie
Liu, Jiaming
Liu, Zhiyuan
Yang, Yicun
Zhang, Linfeng
Zhang, Shanghang
Zhu, Yichen
Xu, Yi
Robotics
Computer Vision and Pattern Recognition
Vision-Language-Action (VLA) models are emerging as a next-generation paradigm for robotics. We introduce dVLA, a diffusion-based VLA that leverages a multimodal chain-of-thought to unify visual perception, language reasoning, and robotic control in a single system. dVLA jointly optimizes perception, language understanding, and action under a single diffusion objective, enabling stronger cross-modal reasoning and better generalization to novel instructions and objects. For practical deployment, we mitigate inference latency by incorporating two acceleration strategies, a prefix attention mask and KV caching, yielding up to around times speedup at test-time inference. We evaluate dVLA in both simulation and the real world: on the LIBERO benchmark, it achieves state-of-the-art performance with a 96.4% average success rate, consistently surpassing both discrete and continuous action policies; on a real Franka robot, it succeeds across a diverse task suite, including a challenging bin-picking task that requires multi-step planning, demonstrating robust real-world performance. Together, these results underscore the promise of unified diffusion frameworks for practical, high-performance VLA robotics.
title dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.25681