SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Haozhan, Zuo, Yuxin, Yu, Jiale, Zhang, Yuhao, Yang, Zhaohui, Zhang, Kaiyan, Zhu, Xuekai, Zhang, Yuchen, Chen, Tianxing, Cui, Ganqu, Wang, Dehui, Luo, Dingxiang, Fan, Yuchen, Sun, Youbang, Zeng, Jia, Pang, Jiangmiao, Zhang, Shanghang, Wang, Yu, Mu, Yao, Zhou, Bowen, Ding, Ning
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912583676592128
author Li, Haozhan
Zuo, Yuxin
Yu, Jiale
Zhang, Yuhao
Yang, Zhaohui
Zhang, Kaiyan
Zhu, Xuekai
Zhang, Yuchen
Chen, Tianxing
Cui, Ganqu
Wang, Dehui
Luo, Dingxiang
Fan, Yuchen
Sun, Youbang
Zeng, Jia
Pang, Jiangmiao
Zhang, Shanghang
Wang, Yu
Mu, Yao
Zhou, Bowen
Ding, Ning
author_facet Li, Haozhan
Zuo, Yuxin
Yu, Jiale
Zhang, Yuhao
Yang, Zhaohui
Zhang, Kaiyan
Zhu, Xuekai
Zhang, Yuchen
Chen, Tianxing
Cui, Ganqu
Wang, Dehui
Luo, Dingxiang
Fan, Yuchen
Sun, Youbang
Zeng, Jia
Pang, Jiangmiao
Zhang, Shanghang
Wang, Yu
Mu, Yao
Zhou, Bowen
Ding, Ning
contents Vision-Language-Action (VLA) models have recently emerged as a powerful paradigm for robotic manipulation. Despite substantial progress enabled by large-scale pretraining and supervised fine-tuning (SFT), these models face two fundamental challenges: (i) the scarcity and high cost of large-scale human-operated robotic trajectories required for SFT scaling, and (ii) limited generalization to tasks involving distribution shift. Recent breakthroughs in Large Reasoning Models (LRMs) demonstrate that reinforcement learning (RL) can dramatically enhance step-by-step reasoning capabilities, raising a natural question: Can RL similarly improve the long-horizon step-by-step action planning of VLA? In this work, we introduce SimpleVLA-RL, an efficient RL framework tailored for VLA models. Building upon veRL, we introduce VLA-specific trajectory sampling, scalable parallelization, multi-environment rendering, and optimized loss computation. When applied to OpenVLA-OFT, SimpleVLA-RL achieves SoTA performance on LIBERO and even outperforms $π_0$ on RoboTwin 1.0\&2.0 with the exploration-enhancing strategies we introduce. SimpleVLA-RL not only reduces dependence on large-scale data and enables robust generalization, but also remarkably surpasses SFT in real-world tasks. Moreover, we identify a novel phenomenon ``pushcut'' during RL training, wherein the policy discovers previously unseen patterns beyond those seen in the previous training process. Github: https://github.com/PRIME-RL/SimpleVLA-RL
format Preprint
id arxiv_https___arxiv_org_abs_2509_09674
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
Li, Haozhan
Zuo, Yuxin
Yu, Jiale
Zhang, Yuhao
Yang, Zhaohui
Zhang, Kaiyan
Zhu, Xuekai
Zhang, Yuchen
Chen, Tianxing
Cui, Ganqu
Wang, Dehui
Luo, Dingxiang
Fan, Yuchen
Sun, Youbang
Zeng, Jia
Pang, Jiangmiao
Zhang, Shanghang
Wang, Yu
Mu, Yao
Zhou, Bowen
Ding, Ning
Robotics
Artificial Intelligence
Computation and Language
Machine Learning
Vision-Language-Action (VLA) models have recently emerged as a powerful paradigm for robotic manipulation. Despite substantial progress enabled by large-scale pretraining and supervised fine-tuning (SFT), these models face two fundamental challenges: (i) the scarcity and high cost of large-scale human-operated robotic trajectories required for SFT scaling, and (ii) limited generalization to tasks involving distribution shift. Recent breakthroughs in Large Reasoning Models (LRMs) demonstrate that reinforcement learning (RL) can dramatically enhance step-by-step reasoning capabilities, raising a natural question: Can RL similarly improve the long-horizon step-by-step action planning of VLA? In this work, we introduce SimpleVLA-RL, an efficient RL framework tailored for VLA models. Building upon veRL, we introduce VLA-specific trajectory sampling, scalable parallelization, multi-environment rendering, and optimized loss computation. When applied to OpenVLA-OFT, SimpleVLA-RL achieves SoTA performance on LIBERO and even outperforms $π_0$ on RoboTwin 1.0\&2.0 with the exploration-enhancing strategies we introduce. SimpleVLA-RL not only reduces dependence on large-scale data and enables robust generalization, but also remarkably surpasses SFT in real-world tasks. Moreover, we identify a novel phenomenon ``pushcut'' during RL training, wherein the policy discovers previously unseen patterns beyond those seen in the previous training process. Github: https://github.com/PRIME-RL/SimpleVLA-RL
title SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
topic Robotics
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2509.09674