VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Hengtao, Ding, Pengxiang, Suo, Runze, Wang, Yihao, Ge, Zirui, Zang, Dongyuan, Yu, Kexian, Sun, Mingyang, Zhang, Hongyin, Wang, Donglin, Su, Weihua
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909818188464128
author Li, Hengtao
Ding, Pengxiang
Suo, Runze
Wang, Yihao
Ge, Zirui
Zang, Dongyuan
Yu, Kexian
Sun, Mingyang
Zhang, Hongyin
Wang, Donglin
Su, Weihua
author_facet Li, Hengtao
Ding, Pengxiang
Suo, Runze
Wang, Yihao
Ge, Zirui
Zang, Dongyuan
Yu, Kexian
Sun, Mingyang
Zhang, Hongyin
Wang, Donglin
Su, Weihua
contents Vision-Language-Action (VLA) models enable embodied decision-making but rely heavily on imitation learning, leading to compounding errors and poor robustness under distribution shift. Reinforcement learning (RL) can mitigate these issues yet typically demands costly real-world interactions or suffers from sim-to-real gaps. We introduce VLA-RFT, a reinforcement fine-tuning framework that leverages a data-driven world model as a controllable simulator. Trained from real interaction data, the simulator predicts future visual observations conditioned on actions, allowing policy rollouts with dense, trajectory-level rewards derived from goal-achieving references. This design delivers an efficient and action-aligned learning signal, drastically lowering sample requirements. With fewer than 400 fine-tuning steps, VLA-RFT surpasses strong supervised baselines and achieves greater efficiency than simulator-based RL. Moreover, it exhibits strong robustness under perturbed conditions, sustaining stable task execution. Our results establish world-model-based RFT as a practical post-training paradigm to enhance the generalization and robustness of VLA models. For more details, please refer to https://vla-rft.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2510_00406
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
Li, Hengtao
Ding, Pengxiang
Suo, Runze
Wang, Yihao
Ge, Zirui
Zang, Dongyuan
Yu, Kexian
Sun, Mingyang
Zhang, Hongyin
Wang, Donglin
Su, Weihua
Robotics
Computer Vision and Pattern Recognition
Vision-Language-Action (VLA) models enable embodied decision-making but rely heavily on imitation learning, leading to compounding errors and poor robustness under distribution shift. Reinforcement learning (RL) can mitigate these issues yet typically demands costly real-world interactions or suffers from sim-to-real gaps. We introduce VLA-RFT, a reinforcement fine-tuning framework that leverages a data-driven world model as a controllable simulator. Trained from real interaction data, the simulator predicts future visual observations conditioned on actions, allowing policy rollouts with dense, trajectory-level rewards derived from goal-achieving references. This design delivers an efficient and action-aligned learning signal, drastically lowering sample requirements. With fewer than 400 fine-tuning steps, VLA-RFT surpasses strong supervised baselines and achieves greater efficiency than simulator-based RL. Moreover, it exhibits strong robustness under perturbed conditions, sustaining stable task execution. Our results establish world-model-based RFT as a practical post-training paradigm to enhance the generalization and robustness of VLA models. For more details, please refer to https://vla-rft.github.io/.
title VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.00406