X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918158761197568 |
|---|---|
| author | Zheng, Jinliang Li, Jianxiong Wang, Zhihao Liu, Dongxiu Kang, Xirui Feng, Yuchun Zheng, Yinan Zou, Jiayin Chen, Yilun Zeng, Jia Zhang, Ya-Qin Pang, Jiangmiao Liu, Jingjing Wang, Tai Zhan, Xianyuan |
| author_facet | Zheng, Jinliang Li, Jianxiong Wang, Zhihao Liu, Dongxiu Kang, Xirui Feng, Yuchun Zheng, Yinan Zou, Jiayin Chen, Yilun Zeng, Jia Zhang, Ya-Qin Pang, Jiangmiao Liu, Jingjing Wang, Tai Zhan, Xianyuan |
| contents | Successful generalist Vision-Language-Action (VLA) models rely on effective training across diverse robotic platforms with large-scale, cross-embodiment, heterogeneous datasets. To facilitate and leverage the heterogeneity in rich, diverse robotic data sources, we propose a novel Soft Prompt approach with minimally added parameters, by infusing prompt learning concepts into cross-embodiment robot learning and introducing separate sets of learnable embeddings for each distinct data source. These embeddings serve as embodiment-specific prompts, which in unity empower VLA models with effective exploitation of varying cross-embodiment features. Our new X-VLA, a neat flow-matching-based VLA architecture, relies exclusively on soft-prompted standard Transformer encoders, enjoying both scalability and simplicity. Evaluated across 6 simulations as well as 3 real-world robots, our 0.9B instantiation-X-VLA-0.9B simultaneously achieves SOTA performance over a sweep of benchmarks, demonstrating superior results on a wide axes of capabilities, from flexible dexterity to quick adaptation across embodiments, environments, and tasks. Website: https://thu-air-dream.github.io/X-VLA/ |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_10274 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model Zheng, Jinliang Li, Jianxiong Wang, Zhihao Liu, Dongxiu Kang, Xirui Feng, Yuchun Zheng, Yinan Zou, Jiayin Chen, Yilun Zeng, Jia Zhang, Ya-Qin Pang, Jiangmiao Liu, Jingjing Wang, Tai Zhan, Xianyuan Robotics Artificial Intelligence Computer Vision and Pattern Recognition Successful generalist Vision-Language-Action (VLA) models rely on effective training across diverse robotic platforms with large-scale, cross-embodiment, heterogeneous datasets. To facilitate and leverage the heterogeneity in rich, diverse robotic data sources, we propose a novel Soft Prompt approach with minimally added parameters, by infusing prompt learning concepts into cross-embodiment robot learning and introducing separate sets of learnable embeddings for each distinct data source. These embeddings serve as embodiment-specific prompts, which in unity empower VLA models with effective exploitation of varying cross-embodiment features. Our new X-VLA, a neat flow-matching-based VLA architecture, relies exclusively on soft-prompted standard Transformer encoders, enjoying both scalability and simplicity. Evaluated across 6 simulations as well as 3 real-world robots, our 0.9B instantiation-X-VLA-0.9B simultaneously achieves SOTA performance over a sweep of benchmarks, demonstrating superior results on a wide axes of capabilities, from flexible dexterity to quick adaptation across embodiments, environments, and tasks. Website: https://thu-air-dream.github.io/X-VLA/ |
| title | X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model |
| topic | Robotics Artificial Intelligence Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2510.10274 |