DeMaVLA: A Vision-Language-Action Foundation Model for Generalizable Deformable Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Su, Taiyi, Zhu, Jian, Wang, Tianjian, He, Youzhang, Huang, Zitai, Zhang, Jianjun, Ma, Chong, Wang, Hanyang, Zhang, Tianjiao, Yin, Munan, Ding, Weihao, Xu, Yi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910272115965952
author Su, Taiyi
Zhu, Jian
Wang, Tianjian
He, Youzhang
Huang, Zitai
Zhang, Jianjun
Ma, Chong
Wang, Hanyang
Zhang, Tianjiao
Yin, Munan
Ding, Weihao
Xu, Yi
author_facet Su, Taiyi
Zhu, Jian
Wang, Tianjian
He, Youzhang
Huang, Zitai
Zhang, Jianjun
Ma, Chong
Wang, Hanyang
Zhang, Tianjiao
Yin, Munan
Ding, Weihao
Xu, Yi
contents Real-world household robots require Vision-Language-Action (VLA) foundation models that can acquire reusable manipulation skills across diverse objects, task conditions, and household environments. Deformable-object folding is a representative challenge, requiring robots to handle clothing items from random initial states across varying categories, geometries, materials, and scenes. However, existing VLA systems commonly train separate policies for different object categories, while naively mixed multi-task training often suffers from task interference and degraded performance. To move beyond category-specific folding policies, we introduce DeMaVLA, a VLA foundation model for generalizable Deformable Manipulation. DeMaVLA adopts a VLM backbone with an action expert and formulates continuous action generation using flow matching. To improve efficiency, the action expert is constructed by pruning every other transformer layer while preserving layer-wise alignment with the VLM backbone, reducing training and inference cost. DeMaVLA is first pre-trained on approximately 5,000 hours of selected real-world dual-arm demonstrations to acquire general manipulation priors. It is then post-trained on mixed folding data that aggregates self-collected demonstrations and corrective trajectories from real-robot failures across multiple folding tasks through a human-in-the-loop Data Aggregation~(DAgger) pipeline. Experiments show that DeMaVLA achieves competitive performance on RoboTwin and strong real-world results on our household folding benchmark. These results highlight the value of scalable real-world data, efficient action generation, and corrective learning for general-purpose VLA policies in deformable-object manipulation.
format Preprint
id arxiv_https___arxiv_org_abs_2605_31286
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DeMaVLA: A Vision-Language-Action Foundation Model for Generalizable Deformable Manipulation
Su, Taiyi
Zhu, Jian
Wang, Tianjian
He, Youzhang
Huang, Zitai
Zhang, Jianjun
Ma, Chong
Wang, Hanyang
Zhang, Tianjiao
Yin, Munan
Ding, Weihao
Xu, Yi
Robotics
Artificial Intelligence
Real-world household robots require Vision-Language-Action (VLA) foundation models that can acquire reusable manipulation skills across diverse objects, task conditions, and household environments. Deformable-object folding is a representative challenge, requiring robots to handle clothing items from random initial states across varying categories, geometries, materials, and scenes. However, existing VLA systems commonly train separate policies for different object categories, while naively mixed multi-task training often suffers from task interference and degraded performance. To move beyond category-specific folding policies, we introduce DeMaVLA, a VLA foundation model for generalizable Deformable Manipulation. DeMaVLA adopts a VLM backbone with an action expert and formulates continuous action generation using flow matching. To improve efficiency, the action expert is constructed by pruning every other transformer layer while preserving layer-wise alignment with the VLM backbone, reducing training and inference cost. DeMaVLA is first pre-trained on approximately 5,000 hours of selected real-world dual-arm demonstrations to acquire general manipulation priors. It is then post-trained on mixed folding data that aggregates self-collected demonstrations and corrective trajectories from real-robot failures across multiple folding tasks through a human-in-the-loop Data Aggregation~(DAgger) pipeline. Experiments show that DeMaVLA achieves competitive performance on RoboTwin and strong real-world results on our household folding benchmark. These results highlight the value of scalable real-world data, efficient action generation, and corrective learning for general-purpose VLA policies in deformable-object manipulation.
title DeMaVLA: A Vision-Language-Action Foundation Model for Generalizable Deformable Manipulation
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2605.31286