Bootstrap Off-policy with World Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhan, Guojian, Wang, Likun, Zhang, Xiangteng, Gao, Jiaxin, Tomizuka, Masayoshi, Li, Shengbo Eben
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917205446230016
author Zhan, Guojian
Wang, Likun
Zhang, Xiangteng
Gao, Jiaxin
Tomizuka, Masayoshi
Li, Shengbo Eben
author_facet Zhan, Guojian
Wang, Likun
Zhang, Xiangteng
Gao, Jiaxin
Tomizuka, Masayoshi
Li, Shengbo Eben
contents Online planning has proven effective in reinforcement learning (RL) for improving sample efficiency and final performance. However, using planning for environment interaction inevitably introduces a divergence between the collected data and the policy's actual behaviors, degrading both model learning and policy improvement. To address this, we propose BOOM (Bootstrap Off-policy with WOrld Model), a framework that tightly integrates planning and off-policy learning through a bootstrap loop: the policy initializes the planner, and the planner refines actions to bootstrap the policy through behavior alignment. This loop is supported by a jointly learned world model, which enables the planner to simulate future trajectories and provides value targets to facilitate policy improvement. The core of BOOM is a likelihood-free alignment loss that bootstraps the policy using the planner's non-parametric action distribution, combined with a soft value-weighted mechanism that prioritizes high-return behaviors and mitigates variability in the planner's action quality within the replay buffer. Experiments on the high-dimensional DeepMind Control Suite and Humanoid-Bench show that BOOM achieves state-of-the-art results in both training stability and final performance. The code is accessible at https://github.com/molumitu/BOOM_MBRL.
format Preprint
id arxiv_https___arxiv_org_abs_2511_00423
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Bootstrap Off-policy with World Model
Zhan, Guojian
Wang, Likun
Zhang, Xiangteng
Gao, Jiaxin
Tomizuka, Masayoshi
Li, Shengbo Eben
Machine Learning
Artificial Intelligence
Robotics
Online planning has proven effective in reinforcement learning (RL) for improving sample efficiency and final performance. However, using planning for environment interaction inevitably introduces a divergence between the collected data and the policy's actual behaviors, degrading both model learning and policy improvement. To address this, we propose BOOM (Bootstrap Off-policy with WOrld Model), a framework that tightly integrates planning and off-policy learning through a bootstrap loop: the policy initializes the planner, and the planner refines actions to bootstrap the policy through behavior alignment. This loop is supported by a jointly learned world model, which enables the planner to simulate future trajectories and provides value targets to facilitate policy improvement. The core of BOOM is a likelihood-free alignment loss that bootstraps the policy using the planner's non-parametric action distribution, combined with a soft value-weighted mechanism that prioritizes high-return behaviors and mitigates variability in the planner's action quality within the replay buffer. Experiments on the high-dimensional DeepMind Control Suite and Humanoid-Bench show that BOOM achieves state-of-the-art results in both training stability and final performance. The code is accessible at https://github.com/molumitu/BOOM_MBRL.
title Bootstrap Off-policy with World Model
topic Machine Learning
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2511.00423