Uncertainty-Based Smooth Policy Regularisation for Reinforcement Learning with Few Demonstrations

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhu, Yujie, Hepburn, Charles A., Thorpe, Matthew, Montana, Giovanni
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908621943603200
author Zhu, Yujie
Hepburn, Charles A.
Thorpe, Matthew
Montana, Giovanni
author_facet Zhu, Yujie
Hepburn, Charles A.
Thorpe, Matthew
Montana, Giovanni
contents In reinforcement learning with sparse rewards, demonstrations can accelerate learning, but determining when to imitate them remains challenging. We propose Smooth Policy Regularisation from Demonstrations (SPReD), a framework that addresses the fundamental question: when should an agent imitate a demonstration versus follow its own policy? SPReD uses ensemble methods to explicitly model Q-value distributions for both demonstration and policy actions, quantifying uncertainty for comparisons. We develop two complementary uncertainty-aware methods: a probabilistic approach estimating the likelihood of demonstration superiority, and an advantage-based approach scaling imitation by statistical significance. Unlike prevailing methods (e.g. Q-filter) that make binary imitation decisions, SPReD applies continuous, uncertainty-proportional regularisation weights, reducing gradient variance during training. Despite its computational simplicity, SPReD achieves remarkable gains in experiments across eight robotics tasks, outperforming existing approaches by up to a factor of 14 in complex tasks while maintaining robustness to demonstration quality and quantity. Our code is available at https://github.com/YujieZhu7/SPReD.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15981
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Uncertainty-Based Smooth Policy Regularisation for Reinforcement Learning with Few Demonstrations
Zhu, Yujie
Hepburn, Charles A.
Thorpe, Matthew
Montana, Giovanni
Machine Learning
Artificial Intelligence
Robotics
In reinforcement learning with sparse rewards, demonstrations can accelerate learning, but determining when to imitate them remains challenging. We propose Smooth Policy Regularisation from Demonstrations (SPReD), a framework that addresses the fundamental question: when should an agent imitate a demonstration versus follow its own policy? SPReD uses ensemble methods to explicitly model Q-value distributions for both demonstration and policy actions, quantifying uncertainty for comparisons. We develop two complementary uncertainty-aware methods: a probabilistic approach estimating the likelihood of demonstration superiority, and an advantage-based approach scaling imitation by statistical significance. Unlike prevailing methods (e.g. Q-filter) that make binary imitation decisions, SPReD applies continuous, uncertainty-proportional regularisation weights, reducing gradient variance during training. Despite its computational simplicity, SPReD achieves remarkable gains in experiments across eight robotics tasks, outperforming existing approaches by up to a factor of 14 in complex tasks while maintaining robustness to demonstration quality and quantity. Our code is available at https://github.com/YujieZhu7/SPReD.
title Uncertainty-Based Smooth Policy Regularisation for Reinforcement Learning with Few Demonstrations
topic Machine Learning
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2509.15981