World4RL: Diffusion World Models for Policy Refinement with Reinforcement Learning for Robotic Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Zhennan, Liu, Kai, Qin, Yuxin, Tian, Shuai, Zheng, Yupeng, Zhou, Mingcai, Yu, Chao, Li, Haoran, Zhao, Dongbin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910060391694336
author Jiang, Zhennan
Liu, Kai
Qin, Yuxin
Tian, Shuai
Zheng, Yupeng
Zhou, Mingcai
Yu, Chao
Li, Haoran
Zhao, Dongbin
author_facet Jiang, Zhennan
Liu, Kai
Qin, Yuxin
Tian, Shuai
Zheng, Yupeng
Zhou, Mingcai
Yu, Chao
Li, Haoran
Zhao, Dongbin
contents Robotic manipulation policies are commonly initialized through imitation learning, but their performance is limited by the scarcity and narrow coverage of expert data. Reinforcement learning can refine polices to alleviate this limitation, yet real-robot training is costly and unsafe, while training in simulators suffers from the sim-to-real gap. Recent advances in generative models have demonstrated remarkable capabilities in real-world simulation, with diffusion models in particular excelling at generation. This raises the question of how diffusion model-based world models can be combined to enhance pre-trained policies in robotic manipulation. In this work, we propose World4RL, a framework that employs diffusion-based world models as high-fidelity simulators to refine pre-trained policies entirely in imagined environments for robotic manipulation. Unlike prior works that primarily employ world models for planning, our framework enables direct end-to-end policy optimization. World4RL is designed around two principles: pre-training a diffusion world model that captures diverse dynamics on multi-task datasets and refining policies entirely within a frozen world model to avoid online real-world interactions. We further design a two-hot action encoding scheme tailored for robotic manipulation and adopt diffusion backbones to improve modeling fidelity. Extensive simulation and real-world experiments demonstrate that World4RL provides high-fidelity environment modeling and enables consistent policy refinement, yielding significantly higher success rates compared to imitation learning and other baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2509_19080
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle World4RL: Diffusion World Models for Policy Refinement with Reinforcement Learning for Robotic Manipulation
Jiang, Zhennan
Liu, Kai
Qin, Yuxin
Tian, Shuai
Zheng, Yupeng
Zhou, Mingcai
Yu, Chao
Li, Haoran
Zhao, Dongbin
Robotics
Artificial Intelligence
Robotic manipulation policies are commonly initialized through imitation learning, but their performance is limited by the scarcity and narrow coverage of expert data. Reinforcement learning can refine polices to alleviate this limitation, yet real-robot training is costly and unsafe, while training in simulators suffers from the sim-to-real gap. Recent advances in generative models have demonstrated remarkable capabilities in real-world simulation, with diffusion models in particular excelling at generation. This raises the question of how diffusion model-based world models can be combined to enhance pre-trained policies in robotic manipulation. In this work, we propose World4RL, a framework that employs diffusion-based world models as high-fidelity simulators to refine pre-trained policies entirely in imagined environments for robotic manipulation. Unlike prior works that primarily employ world models for planning, our framework enables direct end-to-end policy optimization. World4RL is designed around two principles: pre-training a diffusion world model that captures diverse dynamics on multi-task datasets and refining policies entirely within a frozen world model to avoid online real-world interactions. We further design a two-hot action encoding scheme tailored for robotic manipulation and adopt diffusion backbones to improve modeling fidelity. Extensive simulation and real-world experiments demonstrate that World4RL provides high-fidelity environment modeling and enables consistent policy refinement, yielding significantly higher success rates compared to imitation learning and other baselines.
title World4RL: Diffusion World Models for Policy Refinement with Reinforcement Learning for Robotic Manipulation
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2509.19080