How Ensembles of Distilled Policies Improve Generalisation in Reinforcement Learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Weltevrede, Max, Zanger, Moritz A., Spaan, Matthijs T. J., Böhmer, Wendelin
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918167624810496
author Weltevrede, Max
Zanger, Moritz A.
Spaan, Matthijs T. J.
Böhmer, Wendelin
author_facet Weltevrede, Max
Zanger, Moritz A.
Spaan, Matthijs T. J.
Böhmer, Wendelin
contents In the zero-shot policy transfer setting in reinforcement learning, the goal is to train an agent on a fixed set of training environments so that it can generalise to similar, but unseen, testing environments. Previous work has shown that policy distillation after training can sometimes produce a policy that outperforms the original in the testing environments. However, it is not yet entirely clear why that is, or what data should be used to distil the policy. In this paper, we prove, under certain assumptions, a generalisation bound for policy distillation after training. The theory provides two practical insights: for improved generalisation, you should 1) train an ensemble of distilled policies, and 2) distil it on as much data from the training environments as possible. We empirically verify that these insights hold in more general settings, when the assumptions required for the theory no longer hold. Finally, we demonstrate that an ensemble of policies distilled on a diverse dataset can generalise significantly better than the original agent.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16581
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle How Ensembles of Distilled Policies Improve Generalisation in Reinforcement Learning
Weltevrede, Max
Zanger, Moritz A.
Spaan, Matthijs T. J.
Böhmer, Wendelin
Machine Learning
Artificial Intelligence
In the zero-shot policy transfer setting in reinforcement learning, the goal is to train an agent on a fixed set of training environments so that it can generalise to similar, but unseen, testing environments. Previous work has shown that policy distillation after training can sometimes produce a policy that outperforms the original in the testing environments. However, it is not yet entirely clear why that is, or what data should be used to distil the policy. In this paper, we prove, under certain assumptions, a generalisation bound for policy distillation after training. The theory provides two practical insights: for improved generalisation, you should 1) train an ensemble of distilled policies, and 2) distil it on as much data from the training environments as possible. We empirically verify that these insights hold in more general settings, when the assumptions required for the theory no longer hold. Finally, we demonstrate that an ensemble of policies distilled on a diverse dataset can generalise significantly better than the original agent.
title How Ensembles of Distilled Policies Improve Generalisation in Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2505.16581