Reinforcement Learning via Implicit Imitation Guidance

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Dong, Perry, Lessing, Alec M., Chen, Annie S., Finn, Chelsea
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913885135568896
author Dong, Perry
Lessing, Alec M.
Chen, Annie S.
Finn, Chelsea
author_facet Dong, Perry
Lessing, Alec M.
Chen, Annie S.
Finn, Chelsea
contents We study the problem of sample efficient reinforcement learning, where prior data such as demonstrations are provided for initialization in lieu of a dense reward signal. A natural approach is to incorporate an imitation learning objective, either as regularization during training or to acquire a reference policy. However, imitation learning objectives can ultimately degrade long-term performance, as it does not directly align with reward maximization. In this work, we propose to use prior data solely for guiding exploration via noise added to the policy, sidestepping the need for explicit behavior cloning constraints. The key insight in our framework, Data-Guided Noise (DGN), is that demonstrations are most useful for identifying which actions should be explored, rather than forcing the policy to take certain actions. Our approach achieves up to 2-3x improvement over prior reinforcement learning from offline data methods across seven simulated continuous control tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_07505
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reinforcement Learning via Implicit Imitation Guidance
Dong, Perry
Lessing, Alec M.
Chen, Annie S.
Finn, Chelsea
Machine Learning
Artificial Intelligence
We study the problem of sample efficient reinforcement learning, where prior data such as demonstrations are provided for initialization in lieu of a dense reward signal. A natural approach is to incorporate an imitation learning objective, either as regularization during training or to acquire a reference policy. However, imitation learning objectives can ultimately degrade long-term performance, as it does not directly align with reward maximization. In this work, we propose to use prior data solely for guiding exploration via noise added to the policy, sidestepping the need for explicit behavior cloning constraints. The key insight in our framework, Data-Guided Noise (DGN), is that demonstrations are most useful for identifying which actions should be explored, rather than forcing the policy to take certain actions. Our approach achieves up to 2-3x improvement over prior reinforcement learning from offline data methods across seven simulated continuous control tasks.
title Reinforcement Learning via Implicit Imitation Guidance
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.07505