Residual Off-Policy RL for Finetuning Behavior Cloning Policies

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ankile, Lars, Jiang, Zhenyu, Duan, Rocky, Shi, Guanya, Abbeel, Pieter, Nagabandi, Anusha
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915514244136960
author Ankile, Lars
Jiang, Zhenyu
Duan, Rocky
Shi, Guanya
Abbeel, Pieter
Nagabandi, Anusha
author_facet Ankile, Lars
Jiang, Zhenyu
Duan, Rocky
Shi, Guanya
Abbeel, Pieter
Nagabandi, Anusha
contents Recent advances in behavior cloning (BC) have enabled impressive visuomotor control policies. However, these approaches are limited by the quality of human demonstrations, the manual effort required for data collection, and the diminishing returns from offline data. In comparison, reinforcement learning (RL) trains an agent through autonomous interaction with the environment and has shown remarkable success in various domains. Still, training RL policies directly on real-world robots remains challenging due to sample inefficiency, safety concerns, and the difficulty of learning from sparse rewards for long-horizon tasks, especially for high-degree-of-freedom (DoF) systems. We present a recipe that combines the benefits of BC and RL through a residual learning framework. Our approach leverages BC policies as black-box bases and learns lightweight per-step residual corrections via sample-efficient off-policy RL. We demonstrate that our method requires only sparse binary reward signals and can effectively improve manipulation policies on high-degree-of-freedom (DoF) systems in both simulation and the real world. In particular, we demonstrate, to the best of our knowledge, the first successful real-world RL training on a humanoid robot with dexterous hands. Our results demonstrate state-of-the-art performance in various vision-based tasks, pointing towards a practical pathway for deploying RL in the real world.
format Preprint
id arxiv_https___arxiv_org_abs_2509_19301
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Residual Off-Policy RL for Finetuning Behavior Cloning Policies
Ankile, Lars
Jiang, Zhenyu
Duan, Rocky
Shi, Guanya
Abbeel, Pieter
Nagabandi, Anusha
Robotics
Machine Learning
Recent advances in behavior cloning (BC) have enabled impressive visuomotor control policies. However, these approaches are limited by the quality of human demonstrations, the manual effort required for data collection, and the diminishing returns from offline data. In comparison, reinforcement learning (RL) trains an agent through autonomous interaction with the environment and has shown remarkable success in various domains. Still, training RL policies directly on real-world robots remains challenging due to sample inefficiency, safety concerns, and the difficulty of learning from sparse rewards for long-horizon tasks, especially for high-degree-of-freedom (DoF) systems. We present a recipe that combines the benefits of BC and RL through a residual learning framework. Our approach leverages BC policies as black-box bases and learns lightweight per-step residual corrections via sample-efficient off-policy RL. We demonstrate that our method requires only sparse binary reward signals and can effectively improve manipulation policies on high-degree-of-freedom (DoF) systems in both simulation and the real world. In particular, we demonstrate, to the best of our knowledge, the first successful real-world RL training on a humanoid robot with dexterous hands. Our results demonstrate state-of-the-art performance in various vision-based tasks, pointing towards a practical pathway for deploying RL in the real world.
title Residual Off-Policy RL for Finetuning Behavior Cloning Policies
topic Robotics
Machine Learning
url https://arxiv.org/abs/2509.19301