Saved in:
Bibliographic Details
Main Authors: Yang, Jiazhi, Lin, Kunyang, Li, Jinwei, Zhang, Wencong, Lin, Tianwei, Wu, Longyan, Su, Zhizhong, Zhao, Hao, Zhang, Ya-Qin, Chen, Li, Luo, Ping, Yue, Xiangyu, Li, Hongyang
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.11075
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908996655382528
author Yang, Jiazhi
Lin, Kunyang
Li, Jinwei
Zhang, Wencong
Lin, Tianwei
Wu, Longyan
Su, Zhizhong
Zhao, Hao
Zhang, Ya-Qin
Chen, Li
Luo, Ping
Yue, Xiangyu
Li, Hongyang
author_facet Yang, Jiazhi
Lin, Kunyang
Li, Jinwei
Zhang, Wencong
Lin, Tianwei
Wu, Longyan
Su, Zhizhong
Zhao, Hao
Zhang, Ya-Qin
Chen, Li
Luo, Ping
Yue, Xiangyu
Li, Hongyang
contents Despite the sustained scaling on model capacity and data acquisition, Vision-Language-Action (VLA) models remain brittle in contact-rich and dynamic manipulation tasks, where minor execution deviations can compound into failures. While reinforcement learning (RL) offers a principled path to robustness, on-policy RL in the physical world is constrained by safety risk, hardware cost, and environment reset. To bridge this gap, we present RISE, a scalable framework of robotic reinforcement learning via imagination. At its core is a Compositional World Model that (i) predicts multi-view future via a controllable dynamics model, and (ii) evaluates imagined outcomes with a progress value model, producing informative advantages for the policy improvement. Such compositional design allows state and value to be tailored by best-suited yet distinct architectures and objectives. These components are integrated into a closed-loop self-improving pipeline that continuously generates imaginary rollouts, estimates advantages, and updates the policy in imaginary space without costly physical interaction. Across three challenging real-world tasks, RISE yields significant improvement over prior art, with more than +35% absolute performance increase in dynamic brick sorting, +45% for backpack packing, and +35% for box closing, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2602_11075
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RISE: Self-Improving Robot Policy with Compositional World Model
Yang, Jiazhi
Lin, Kunyang
Li, Jinwei
Zhang, Wencong
Lin, Tianwei
Wu, Longyan
Su, Zhizhong
Zhao, Hao
Zhang, Ya-Qin
Chen, Li
Luo, Ping
Yue, Xiangyu
Li, Hongyang
Robotics
Despite the sustained scaling on model capacity and data acquisition, Vision-Language-Action (VLA) models remain brittle in contact-rich and dynamic manipulation tasks, where minor execution deviations can compound into failures. While reinforcement learning (RL) offers a principled path to robustness, on-policy RL in the physical world is constrained by safety risk, hardware cost, and environment reset. To bridge this gap, we present RISE, a scalable framework of robotic reinforcement learning via imagination. At its core is a Compositional World Model that (i) predicts multi-view future via a controllable dynamics model, and (ii) evaluates imagined outcomes with a progress value model, producing informative advantages for the policy improvement. Such compositional design allows state and value to be tailored by best-suited yet distinct architectures and objectives. These components are integrated into a closed-loop self-improving pipeline that continuously generates imaginary rollouts, estimates advantages, and updates the policy in imaginary space without costly physical interaction. Across three challenging real-world tasks, RISE yields significant improvement over prior art, with more than +35% absolute performance increase in dynamic brick sorting, +45% for backpack packing, and +35% for box closing, respectively.
title RISE: Self-Improving Robot Policy with Compositional World Model
topic Robotics
url https://arxiv.org/abs/2602.11075