Steering Your Diffusion Policy with Latent Space Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wagenmaker, Andrew, Nakamoto, Mitsuhiko, Zhang, Yunchu, Park, Seohong, Yagoub, Waleed, Nagabandi, Anusha, Gupta, Abhishek, Levine, Sergey
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909660141846528
author Wagenmaker, Andrew
Nakamoto, Mitsuhiko
Zhang, Yunchu
Park, Seohong
Yagoub, Waleed
Nagabandi, Anusha
Gupta, Abhishek
Levine, Sergey
author_facet Wagenmaker, Andrew
Nakamoto, Mitsuhiko
Zhang, Yunchu
Park, Seohong
Yagoub, Waleed
Nagabandi, Anusha
Gupta, Abhishek
Levine, Sergey
contents Robotic control policies learned from human demonstrations have achieved impressive results in many real-world applications. However, in scenarios where initial performance is not satisfactory, as is often the case in novel open-world settings, such behavioral cloning (BC)-learned policies typically require collecting additional human demonstrations to further improve their behavior -- an expensive and time-consuming process. In contrast, reinforcement learning (RL) holds the promise of enabling autonomous online policy improvement, but often falls short of achieving this due to the large number of samples it typically requires. In this work we take steps towards enabling fast autonomous adaptation of BC-trained policies via efficient real-world RL. Focusing in particular on diffusion policies -- a state-of-the-art BC methodology -- we propose diffusion steering via reinforcement learning (DSRL): adapting the BC policy by running RL over its latent-noise space. We show that DSRL is highly sample efficient, requires only black-box access to the BC policy, and enables effective real-world autonomous policy improvement. Furthermore, DSRL avoids many of the challenges associated with finetuning diffusion policies, obviating the need to modify the weights of the base policy at all. We demonstrate DSRL on simulated benchmarks, real-world robotic tasks, and for adapting pretrained generalist policies, illustrating its sample efficiency and effective performance at real-world policy improvement.
format Preprint
id arxiv_https___arxiv_org_abs_2506_15799
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Steering Your Diffusion Policy with Latent Space Reinforcement Learning
Wagenmaker, Andrew
Nakamoto, Mitsuhiko
Zhang, Yunchu
Park, Seohong
Yagoub, Waleed
Nagabandi, Anusha
Gupta, Abhishek
Levine, Sergey
Robotics
Machine Learning
Robotic control policies learned from human demonstrations have achieved impressive results in many real-world applications. However, in scenarios where initial performance is not satisfactory, as is often the case in novel open-world settings, such behavioral cloning (BC)-learned policies typically require collecting additional human demonstrations to further improve their behavior -- an expensive and time-consuming process. In contrast, reinforcement learning (RL) holds the promise of enabling autonomous online policy improvement, but often falls short of achieving this due to the large number of samples it typically requires. In this work we take steps towards enabling fast autonomous adaptation of BC-trained policies via efficient real-world RL. Focusing in particular on diffusion policies -- a state-of-the-art BC methodology -- we propose diffusion steering via reinforcement learning (DSRL): adapting the BC policy by running RL over its latent-noise space. We show that DSRL is highly sample efficient, requires only black-box access to the BC policy, and enables effective real-world autonomous policy improvement. Furthermore, DSRL avoids many of the challenges associated with finetuning diffusion policies, obviating the need to modify the weights of the base policy at all. We demonstrate DSRL on simulated benchmarks, real-world robotic tasks, and for adapting pretrained generalist policies, illustrating its sample efficiency and effective performance at real-world policy improvement.
title Steering Your Diffusion Policy with Latent Space Reinforcement Learning
topic Robotics
Machine Learning
url https://arxiv.org/abs/2506.15799