Rethinking Visual-Language-Action Model Scaling: Alignment, Mixture, and Regularization

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Ye, Zheng, Sipeng, Luo, Hao, Zhang, Wanpeng, Yuan, Haoqi, Xu, Chaoyi, Xu, Haiweng, Feng, Yicheng, Yu, Mingyang, Kang, Zhiyu, Lu, Zongqing, Jin, Qin
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908826098204672
author Wang, Ye
Zheng, Sipeng
Luo, Hao
Zhang, Wanpeng
Yuan, Haoqi
Xu, Chaoyi
Xu, Haiweng
Feng, Yicheng
Yu, Mingyang
Kang, Zhiyu
Lu, Zongqing
Jin, Qin
author_facet Wang, Ye
Zheng, Sipeng
Luo, Hao
Zhang, Wanpeng
Yuan, Haoqi
Xu, Chaoyi
Xu, Haiweng
Feng, Yicheng
Yu, Mingyang
Kang, Zhiyu
Lu, Zongqing
Jin, Qin
contents While Vision-Language-Action (VLA) models show strong promise for generalist robot control, it remains unclear whether -- and under what conditions -- the standard "scale data" recipe translates to robotics, where training data is inherently heterogeneous across embodiments, sensors, and action spaces. We present a systematic, controlled study of VLA scaling that revisits core training choices for pretraining across diverse robots. Using a representative VLA framework that combines a vision-language backbone with flow-matching, we ablate key design decisions under matched conditions and evaluate in extensive simulation and real-robot experiments. To improve the reliability of real-world results, we introduce a Grouped Blind Ensemble protocol that blinds operators to model identity and separates policy execution from outcome judgment, reducing experimenter bias. Our analysis targets three dimensions of VLA scaling. (1) Physical alignment: we show that a unified end-effector (EEF)-relative action representation is critical for robust cross-embodiment transfer. (2) Embodiment mixture: we find that naively pooling heterogeneous robot datasets often induces negative transfer rather than gains, underscoring the fragility of indiscriminate data scaling. (3) Training regularization: we observe that intuitive strategies, such as sensory dropout and multi-stage fine-tuning, do not consistently improve performance at scale. Together, this study challenge some common assumptions about embodied scaling and provide practical guidance for training large-scale VLA policies from diverse robotic data. Project website: https://research.beingbeyond.com/rethink_vla
format Preprint
id arxiv_https___arxiv_org_abs_2602_09722
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Rethinking Visual-Language-Action Model Scaling: Alignment, Mixture, and Regularization
Wang, Ye
Zheng, Sipeng
Luo, Hao
Zhang, Wanpeng
Yuan, Haoqi
Xu, Chaoyi
Xu, Haiweng
Feng, Yicheng
Yu, Mingyang
Kang, Zhiyu
Lu, Zongqing
Jin, Qin
Robotics
While Vision-Language-Action (VLA) models show strong promise for generalist robot control, it remains unclear whether -- and under what conditions -- the standard "scale data" recipe translates to robotics, where training data is inherently heterogeneous across embodiments, sensors, and action spaces. We present a systematic, controlled study of VLA scaling that revisits core training choices for pretraining across diverse robots. Using a representative VLA framework that combines a vision-language backbone with flow-matching, we ablate key design decisions under matched conditions and evaluate in extensive simulation and real-robot experiments. To improve the reliability of real-world results, we introduce a Grouped Blind Ensemble protocol that blinds operators to model identity and separates policy execution from outcome judgment, reducing experimenter bias. Our analysis targets three dimensions of VLA scaling. (1) Physical alignment: we show that a unified end-effector (EEF)-relative action representation is critical for robust cross-embodiment transfer. (2) Embodiment mixture: we find that naively pooling heterogeneous robot datasets often induces negative transfer rather than gains, underscoring the fragility of indiscriminate data scaling. (3) Training regularization: we observe that intuitive strategies, such as sensory dropout and multi-stage fine-tuning, do not consistently improve performance at scale. Together, this study challenge some common assumptions about embodied scaling and provide practical guidance for training large-scale VLA policies from diverse robotic data. Project website: https://research.beingbeyond.com/rethink_vla
title Rethinking Visual-Language-Action Model Scaling: Alignment, Mixture, and Regularization
topic Robotics
url https://arxiv.org/abs/2602.09722