Self-Improving Vision-Language-Action Models with Data Generation via Residual RL
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866914127425830912 |
|---|---|
| author | Xiao, Wenli Lin, Haotian Peng, Andy Xue, Haoru He, Tairan Xie, Yuqi Hu, Fengyuan Wu, Jimmy Luo, Zhengyi Fan, Linxi "Jim" Shi, Guanya Zhu, Yuke |
| author_facet | Xiao, Wenli Lin, Haotian Peng, Andy Xue, Haoru He, Tairan Xie, Yuqi Hu, Fengyuan Wu, Jimmy Luo, Zhengyi Fan, Linxi "Jim" Shi, Guanya Zhu, Yuke |
| contents | Supervised fine-tuning (SFT) has become the de facto post-training strategy for large vision-language-action (VLA) models, but its reliance on costly human demonstrations limits scalability and generalization. We propose Probe, Learn, Distill (PLD), a three-stage plug-and-play framework that improves VLAs through residual reinforcement learning (RL) and distribution-aware data collection. In Stage 1, we train lightweight residual actors to probe failure regions of the VLA generalist. In Stage 2, we use a hybrid rollout scheme that aligns collected trajectories with the generalist's deployment distribution while capturing recovery behaviors. In Stage 3, we distill the curated trajectories back into the generalist with standard SFT. PLD achieves near-saturated 99% task success on LIBERO, over 50% gains in SimplerEnv, and 100% success on real-world Franka and YAM arm manipulation tasks. Ablations show that residual probing and distribution-aware replay are key to collecting deployment-aligned data that improves both seen and unseen tasks, offering a scalable path toward self-improving VLA models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_00091 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Self-Improving Vision-Language-Action Models with Data Generation via Residual RL Xiao, Wenli Lin, Haotian Peng, Andy Xue, Haoru He, Tairan Xie, Yuqi Hu, Fengyuan Wu, Jimmy Luo, Zhengyi Fan, Linxi "Jim" Shi, Guanya Zhu, Yuke Computer Vision and Pattern Recognition Robotics Supervised fine-tuning (SFT) has become the de facto post-training strategy for large vision-language-action (VLA) models, but its reliance on costly human demonstrations limits scalability and generalization. We propose Probe, Learn, Distill (PLD), a three-stage plug-and-play framework that improves VLAs through residual reinforcement learning (RL) and distribution-aware data collection. In Stage 1, we train lightweight residual actors to probe failure regions of the VLA generalist. In Stage 2, we use a hybrid rollout scheme that aligns collected trajectories with the generalist's deployment distribution while capturing recovery behaviors. In Stage 3, we distill the curated trajectories back into the generalist with standard SFT. PLD achieves near-saturated 99% task success on LIBERO, over 50% gains in SimplerEnv, and 100% success on real-world Franka and YAM arm manipulation tasks. Ablations show that residual probing and distribution-aware replay are key to collecting deployment-aligned data that improves both seen and unseen tasks, offering a scalable path toward self-improving VLA models. |
| title | Self-Improving Vision-Language-Action Models with Data Generation via Residual RL |
| topic | Computer Vision and Pattern Recognition Robotics |
| url | https://arxiv.org/abs/2511.00091 |