RynnVLA-002: A Unified Vision-Language-Action and World Model

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cen, Jun, Huang, Siteng, Yuan, Yuqian, Li, Kehan, Yuan, Hangjie, Yu, Chaohui, Hou, Bohan, Jiang, Yuming, Guo, Jiayan, Li, Xin, Luo, Hao, Wang, Fan, Zhao, Deli, Chen, Hao
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913174569091072
author Cen, Jun
Huang, Siteng
Yuan, Yuqian
Li, Kehan
Yuan, Hangjie
Yu, Chaohui
Hou, Bohan
Jiang, Yuming
Guo, Jiayan
Li, Xin
Luo, Hao
Wang, Fan
Zhao, Deli
Chen, Hao
author_facet Cen, Jun
Huang, Siteng
Yuan, Yuqian
Li, Kehan
Yuan, Hangjie
Yu, Chaohui
Hou, Bohan
Jiang, Yuming
Guo, Jiayan
Li, Xin
Luo, Hao
Wang, Fan
Zhao, Deli
Chen, Hao
contents We introduce RynnVLA-002, a unified Vision-Language-Action (VLA) and world model. The world model leverages action and visual inputs to predict future image states, learning the underlying physics of the environment to refine action generation. Conversely, the VLA model produces subsequent actions from image observations, enhancing visual understanding and supporting the world model's image generation. The unified framework of RynnVLA-002 enables joint learning of environmental dynamics and action planning. Our experiments show that RynnVLA-002 surpasses individual VLA and world models, demonstrating their mutual enhancement. We evaluate RynnVLA-002 in both simulation and real-world robot tasks. RynnVLA-002 achieves 97.4% success rate on the LIBERO simulation benchmark without pretraining, while in real-world LeRobot experiments, its integrated world model boosts the overall success rate by 50%.
format Preprint
id arxiv_https___arxiv_org_abs_2511_17502
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RynnVLA-002: A Unified Vision-Language-Action and World Model
Cen, Jun
Huang, Siteng
Yuan, Yuqian
Li, Kehan
Yuan, Hangjie
Yu, Chaohui
Hou, Bohan
Jiang, Yuming
Guo, Jiayan
Li, Xin
Luo, Hao
Wang, Fan
Zhao, Deli
Chen, Hao
Robotics
We introduce RynnVLA-002, a unified Vision-Language-Action (VLA) and world model. The world model leverages action and visual inputs to predict future image states, learning the underlying physics of the environment to refine action generation. Conversely, the VLA model produces subsequent actions from image observations, enhancing visual understanding and supporting the world model's image generation. The unified framework of RynnVLA-002 enables joint learning of environmental dynamics and action planning. Our experiments show that RynnVLA-002 surpasses individual VLA and world models, demonstrating their mutual enhancement. We evaluate RynnVLA-002 in both simulation and real-world robot tasks. RynnVLA-002 achieves 97.4% success rate on the LIBERO simulation benchmark without pretraining, while in real-world LeRobot experiments, its integrated world model boosts the overall success rate by 50%.
title RynnVLA-002: A Unified Vision-Language-Action and World Model
topic Robotics
url https://arxiv.org/abs/2511.17502