Coherent Off-Policy Improvement of Large Behavior Models with Learned Rewards

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Scherer, Christian, Watson, Joe, Gruner, Theo, Palenicek, Daniel, Posner, Ingmar, Peters, Jan
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918535172718592
author Scherer, Christian
Watson, Joe
Gruner, Theo
Palenicek, Daniel
Posner, Ingmar
Peters, Jan
author_facet Scherer, Christian
Watson, Joe
Gruner, Theo
Palenicek, Daniel
Posner, Ingmar
Peters, Jan
contents Distilling expert demonstration data into large generative models using behavioral cloning is a scalable approach to learning capable policies for robotic control, particularly for dexterous manipulation. Reinforcement learning (RL) can be used as a means to finetune these policies further using additional experience. An open question is whether RL is more sample-efficient than collecting more human demonstrations. Prior work has finetuned large pretrained policies in a scalable fashion by applying RL to a smaller residual policy that corrects the pretrained model. However, for the typical sparse reward tasks, RL algorithms can struggle to optimize the behavior in a sample-efficient manner. We explore inverse reinforcement learning, where a dense reward function is learned from expert demonstrations, potentially reducing the challenge of RL finetuning. We specifically consider coherent imitation learning, an IRL method that facilitates improvement of the BC policy through using a specific reward formulation with theoretical guarantees. We show that our IRL method maintains or improves the performance of pi-0.5 on all six sparse manipulation tasks and achieves a $\geq 90\%$ success rate on five out of six complex manipulation tasks, outperforming RL-based baselines using sparse rewards. By ensuring our initial pretrained finetuning policy is optimal for our initial reward and critic, our method circumvents the initial drop commonly seen in RL finetuning and enables faster improvement.
format Preprint
id arxiv_https___arxiv_org_abs_2606_02194
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Coherent Off-Policy Improvement of Large Behavior Models with Learned Rewards
Scherer, Christian
Watson, Joe
Gruner, Theo
Palenicek, Daniel
Posner, Ingmar
Peters, Jan
Machine Learning
Distilling expert demonstration data into large generative models using behavioral cloning is a scalable approach to learning capable policies for robotic control, particularly for dexterous manipulation. Reinforcement learning (RL) can be used as a means to finetune these policies further using additional experience. An open question is whether RL is more sample-efficient than collecting more human demonstrations. Prior work has finetuned large pretrained policies in a scalable fashion by applying RL to a smaller residual policy that corrects the pretrained model. However, for the typical sparse reward tasks, RL algorithms can struggle to optimize the behavior in a sample-efficient manner. We explore inverse reinforcement learning, where a dense reward function is learned from expert demonstrations, potentially reducing the challenge of RL finetuning. We specifically consider coherent imitation learning, an IRL method that facilitates improvement of the BC policy through using a specific reward formulation with theoretical guarantees. We show that our IRL method maintains or improves the performance of pi-0.5 on all six sparse manipulation tasks and achieves a $\geq 90\%$ success rate on five out of six complex manipulation tasks, outperforming RL-based baselines using sparse rewards. By ensuring our initial pretrained finetuning policy is optimal for our initial reward and critic, our method circumvents the initial drop commonly seen in RL finetuning and enables faster improvement.
title Coherent Off-Policy Improvement of Large Behavior Models with Learned Rewards
topic Machine Learning
url https://arxiv.org/abs/2606.02194