X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Jinliang, Li, Jianxiong, Wang, Zhihao, Liu, Dongxiu, Kang, Xirui, Feng, Yuchun, Zheng, Yinan, Zou, Jiayin, Chen, Yilun, Zeng, Jia, Zhang, Ya-Qin, Pang, Jiangmiao, Liu, Jingjing, Wang, Tai, Zhan, Xianyuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918158761197568
author Zheng, Jinliang
Li, Jianxiong
Wang, Zhihao
Liu, Dongxiu
Kang, Xirui
Feng, Yuchun
Zheng, Yinan
Zou, Jiayin
Chen, Yilun
Zeng, Jia
Zhang, Ya-Qin
Pang, Jiangmiao
Liu, Jingjing
Wang, Tai
Zhan, Xianyuan
author_facet Zheng, Jinliang
Li, Jianxiong
Wang, Zhihao
Liu, Dongxiu
Kang, Xirui
Feng, Yuchun
Zheng, Yinan
Zou, Jiayin
Chen, Yilun
Zeng, Jia
Zhang, Ya-Qin
Pang, Jiangmiao
Liu, Jingjing
Wang, Tai
Zhan, Xianyuan
contents Successful generalist Vision-Language-Action (VLA) models rely on effective training across diverse robotic platforms with large-scale, cross-embodiment, heterogeneous datasets. To facilitate and leverage the heterogeneity in rich, diverse robotic data sources, we propose a novel Soft Prompt approach with minimally added parameters, by infusing prompt learning concepts into cross-embodiment robot learning and introducing separate sets of learnable embeddings for each distinct data source. These embeddings serve as embodiment-specific prompts, which in unity empower VLA models with effective exploitation of varying cross-embodiment features. Our new X-VLA, a neat flow-matching-based VLA architecture, relies exclusively on soft-prompted standard Transformer encoders, enjoying both scalability and simplicity. Evaluated across 6 simulations as well as 3 real-world robots, our 0.9B instantiation-X-VLA-0.9B simultaneously achieves SOTA performance over a sweep of benchmarks, demonstrating superior results on a wide axes of capabilities, from flexible dexterity to quick adaptation across embodiments, environments, and tasks. Website: https://thu-air-dream.github.io/X-VLA/
format Preprint
id arxiv_https___arxiv_org_abs_2510_10274
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
Zheng, Jinliang
Li, Jianxiong
Wang, Zhihao
Liu, Dongxiu
Kang, Xirui
Feng, Yuchun
Zheng, Yinan
Zou, Jiayin
Chen, Yilun
Zeng, Jia
Zhang, Ya-Qin
Pang, Jiangmiao
Liu, Jingjing
Wang, Tai
Zhan, Xianyuan
Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Successful generalist Vision-Language-Action (VLA) models rely on effective training across diverse robotic platforms with large-scale, cross-embodiment, heterogeneous datasets. To facilitate and leverage the heterogeneity in rich, diverse robotic data sources, we propose a novel Soft Prompt approach with minimally added parameters, by infusing prompt learning concepts into cross-embodiment robot learning and introducing separate sets of learnable embeddings for each distinct data source. These embeddings serve as embodiment-specific prompts, which in unity empower VLA models with effective exploitation of varying cross-embodiment features. Our new X-VLA, a neat flow-matching-based VLA architecture, relies exclusively on soft-prompted standard Transformer encoders, enjoying both scalability and simplicity. Evaluated across 6 simulations as well as 3 real-world robots, our 0.9B instantiation-X-VLA-0.9B simultaneously achieves SOTA performance over a sweep of benchmarks, demonstrating superior results on a wide axes of capabilities, from flexible dexterity to quick adaptation across embodiments, environments, and tasks. Website: https://thu-air-dream.github.io/X-VLA/
title X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
topic Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.10274