Self-Improving Vision-Language-Action Models with Data Generation via Residual RL

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xiao, Wenli, Lin, Haotian, Peng, Andy, Xue, Haoru, He, Tairan, Xie, Yuqi, Hu, Fengyuan, Wu, Jimmy, Luo, Zhengyi, Fan, Linxi "Jim", Shi, Guanya, Zhu, Yuke
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914127425830912
author Xiao, Wenli
Lin, Haotian
Peng, Andy
Xue, Haoru
He, Tairan
Xie, Yuqi
Hu, Fengyuan
Wu, Jimmy
Luo, Zhengyi
Fan, Linxi "Jim"
Shi, Guanya
Zhu, Yuke
author_facet Xiao, Wenli
Lin, Haotian
Peng, Andy
Xue, Haoru
He, Tairan
Xie, Yuqi
Hu, Fengyuan
Wu, Jimmy
Luo, Zhengyi
Fan, Linxi "Jim"
Shi, Guanya
Zhu, Yuke
contents Supervised fine-tuning (SFT) has become the de facto post-training strategy for large vision-language-action (VLA) models, but its reliance on costly human demonstrations limits scalability and generalization. We propose Probe, Learn, Distill (PLD), a three-stage plug-and-play framework that improves VLAs through residual reinforcement learning (RL) and distribution-aware data collection. In Stage 1, we train lightweight residual actors to probe failure regions of the VLA generalist. In Stage 2, we use a hybrid rollout scheme that aligns collected trajectories with the generalist's deployment distribution while capturing recovery behaviors. In Stage 3, we distill the curated trajectories back into the generalist with standard SFT. PLD achieves near-saturated 99% task success on LIBERO, over 50% gains in SimplerEnv, and 100% success on real-world Franka and YAM arm manipulation tasks. Ablations show that residual probing and distribution-aware replay are key to collecting deployment-aligned data that improves both seen and unseen tasks, offering a scalable path toward self-improving VLA models.
format Preprint
id arxiv_https___arxiv_org_abs_2511_00091
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Self-Improving Vision-Language-Action Models with Data Generation via Residual RL
Xiao, Wenli
Lin, Haotian
Peng, Andy
Xue, Haoru
He, Tairan
Xie, Yuqi
Hu, Fengyuan
Wu, Jimmy
Luo, Zhengyi
Fan, Linxi "Jim"
Shi, Guanya
Zhu, Yuke
Computer Vision and Pattern Recognition
Robotics
Supervised fine-tuning (SFT) has become the de facto post-training strategy for large vision-language-action (VLA) models, but its reliance on costly human demonstrations limits scalability and generalization. We propose Probe, Learn, Distill (PLD), a three-stage plug-and-play framework that improves VLAs through residual reinforcement learning (RL) and distribution-aware data collection. In Stage 1, we train lightweight residual actors to probe failure regions of the VLA generalist. In Stage 2, we use a hybrid rollout scheme that aligns collected trajectories with the generalist's deployment distribution while capturing recovery behaviors. In Stage 3, we distill the curated trajectories back into the generalist with standard SFT. PLD achieves near-saturated 99% task success on LIBERO, over 50% gains in SimplerEnv, and 100% success on real-world Franka and YAM arm manipulation tasks. Ablations show that residual probing and distribution-aware replay are key to collecting deployment-aligned data that improves both seen and unseen tasks, offering a scalable path toward self-improving VLA models.
title Self-Improving Vision-Language-Action Models with Data Generation via Residual RL
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2511.00091