SkillVLA: Tackling Combinatorial Diversity in Dual-Arm Manipulation via Skill Reuse
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866912942156414976 |
|---|---|
| author | Zhai, Xuanran Huang, Zekai Wu, Longyan Zhao, Qianyou Yu, Qiaojun Ren, Jieji Hao, Ce Soh, Harold |
| author_facet | Zhai, Xuanran Huang, Zekai Wu, Longyan Zhao, Qianyou Yu, Qiaojun Ren, Jieji Hao, Ce Soh, Harold |
| contents | Recent progress in vision-language-action (VLA) models has demonstrated strong potential for dual-arm manipulation, enabling complex behaviors and generalization to unseen environments. However, mainstream bimanual VLA formulations largely overlook the critical challenge of combinatorial diversity. Different pairings of single-arm behaviors can induce qualitatively distinct task behaviors, yet existing models do not explicitly account for this structure. We argue that effective bimanual VLAs should support skill reuse - the ability to recombine previously learned single-arm skills across novel left-right pairings - thereby avoiding the need to separately learn every possible combination. Current VLA designs entangle skills across arms, preventing such recomposition and limiting scalability. To address this limitation, we propose SkillVLA, a framework explicitly designed to enable skill reuse in dual-arm manipulation. Extensive experiments demonstrate that SkillVLA substantially improves skill composition, increasing overall success rate from 0% to 51%, and achieves strong performance on cooperative and long-horizon tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_03836 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | SkillVLA: Tackling Combinatorial Diversity in Dual-Arm Manipulation via Skill Reuse Zhai, Xuanran Huang, Zekai Wu, Longyan Zhao, Qianyou Yu, Qiaojun Ren, Jieji Hao, Ce Soh, Harold Robotics Recent progress in vision-language-action (VLA) models has demonstrated strong potential for dual-arm manipulation, enabling complex behaviors and generalization to unseen environments. However, mainstream bimanual VLA formulations largely overlook the critical challenge of combinatorial diversity. Different pairings of single-arm behaviors can induce qualitatively distinct task behaviors, yet existing models do not explicitly account for this structure. We argue that effective bimanual VLAs should support skill reuse - the ability to recombine previously learned single-arm skills across novel left-right pairings - thereby avoiding the need to separately learn every possible combination. Current VLA designs entangle skills across arms, preventing such recomposition and limiting scalability. To address this limitation, we propose SkillVLA, a framework explicitly designed to enable skill reuse in dual-arm manipulation. Extensive experiments demonstrate that SkillVLA substantially improves skill composition, increasing overall success rate from 0% to 51%, and achieves strong performance on cooperative and long-horizon tasks. |
| title | SkillVLA: Tackling Combinatorial Diversity in Dual-Arm Manipulation via Skill Reuse |
| topic | Robotics |
| url | https://arxiv.org/abs/2603.03836 |