DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wen, Junjie, Zhu, Yichen, Li, Jinming, Tang, Zhibin, Shen, Chaomin, Feng, Feifei
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918119945011200
author Wen, Junjie
Zhu, Yichen
Li, Jinming
Tang, Zhibin
Shen, Chaomin
Feng, Feifei
author_facet Wen, Junjie
Zhu, Yichen
Li, Jinming
Tang, Zhibin
Shen, Chaomin
Feng, Feifei
contents Enabling robots to perform diverse tasks across varied environments is a central challenge in robot learning. While vision-language-action (VLA) models have shown promise for generalizable robot skills, realizing their full potential requires addressing limitations in action representation and efficient training. Current VLA models often focus on scaling the vision-language model (VLM) component, while the action space representation remains a critical bottleneck. This paper introduces DexVLA, a novel framework designed to enhance the efficiency and generalization capabilities of VLAs for complex, long-horizon tasks across diverse robot embodiments. DexVLA features a novel diffusion-based action expert, scaled to one billion parameters, designed for cross-embodiment learning. A novel embodiment curriculum learning strategy facilitates efficient training: (1) pre-training the diffusion expert that is separable from the VLA on cross-embodiment data, (2) aligning the VLA model to specific embodiments, and (3) post-training for rapid adaptation to new tasks. We conduct comprehensive experiments across multiple embodiments, including single-arm, bimanual, and dexterous hand, demonstrating DexVLA's adaptability to challenging tasks without task-specific adaptation, its ability to learn dexterous skills on novel embodiments with limited data, and its capacity to complete complex, long-horizon tasks using only direct language prompting, such as laundry folding. In all settings, our method demonstrates superior performance compared to state-of-the-art models like Octo, OpenVLA, and Diffusion Policy.
format Preprint
id arxiv_https___arxiv_org_abs_2502_05855
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control
Wen, Junjie
Zhu, Yichen
Li, Jinming
Tang, Zhibin
Shen, Chaomin
Feng, Feifei
Robotics
Computer Vision and Pattern Recognition
Enabling robots to perform diverse tasks across varied environments is a central challenge in robot learning. While vision-language-action (VLA) models have shown promise for generalizable robot skills, realizing their full potential requires addressing limitations in action representation and efficient training. Current VLA models often focus on scaling the vision-language model (VLM) component, while the action space representation remains a critical bottleneck. This paper introduces DexVLA, a novel framework designed to enhance the efficiency and generalization capabilities of VLAs for complex, long-horizon tasks across diverse robot embodiments. DexVLA features a novel diffusion-based action expert, scaled to one billion parameters, designed for cross-embodiment learning. A novel embodiment curriculum learning strategy facilitates efficient training: (1) pre-training the diffusion expert that is separable from the VLA on cross-embodiment data, (2) aligning the VLA model to specific embodiments, and (3) post-training for rapid adaptation to new tasks. We conduct comprehensive experiments across multiple embodiments, including single-arm, bimanual, and dexterous hand, demonstrating DexVLA's adaptability to challenging tasks without task-specific adaptation, its ability to learn dexterous skills on novel embodiments with limited data, and its capacity to complete complex, long-horizon tasks using only direct language prompting, such as laundry folding. In all settings, our method demonstrates superior performance compared to state-of-the-art models like Octo, OpenVLA, and Diffusion Policy.
title DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.05855