How Should We Meta-Learn Reinforcement Learning Algorithms?

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Goldie, Alexander David, Wang, Zilin, Cohen, Jaron, Foerster, Jakob Nicolaus, Whiteson, Shimon
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915488223723520
author Goldie, Alexander David
Wang, Zilin
Cohen, Jaron
Foerster, Jakob Nicolaus
Whiteson, Shimon
author_facet Goldie, Alexander David
Wang, Zilin
Cohen, Jaron
Foerster, Jakob Nicolaus
Whiteson, Shimon
contents The process of meta-learning algorithms from data, instead of relying on manual design, is growing in popularity as a paradigm for improving the performance of machine learning systems. Meta-learning shows particular promise for reinforcement learning (RL), where algorithms are often adapted from supervised or unsupervised learning despite their suboptimality for RL. However, until now there has been a severe lack of comparison between different meta-learning algorithms, such as using evolution to optimise over black-box functions or LLMs to propose code. In this paper, we carry out this empirical comparison of the different approaches when applied to a range of meta-learned algorithms which target different parts of the RL pipeline. In addition to meta-train and meta-test performance, we also investigate factors including the interpretability, sample cost and train time for each meta-learning algorithm. Based on these findings, we propose several guidelines for meta-learning new RL algorithms which will help ensure that future learned algorithms are as performant as possible.
format Preprint
id arxiv_https___arxiv_org_abs_2507_17668
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle How Should We Meta-Learn Reinforcement Learning Algorithms?
Goldie, Alexander David
Wang, Zilin
Cohen, Jaron
Foerster, Jakob Nicolaus
Whiteson, Shimon
Machine Learning
Artificial Intelligence
The process of meta-learning algorithms from data, instead of relying on manual design, is growing in popularity as a paradigm for improving the performance of machine learning systems. Meta-learning shows particular promise for reinforcement learning (RL), where algorithms are often adapted from supervised or unsupervised learning despite their suboptimality for RL. However, until now there has been a severe lack of comparison between different meta-learning algorithms, such as using evolution to optimise over black-box functions or LLMs to propose code. In this paper, we carry out this empirical comparison of the different approaches when applied to a range of meta-learned algorithms which target different parts of the RL pipeline. In addition to meta-train and meta-test performance, we also investigate factors including the interpretability, sample cost and train time for each meta-learning algorithm. Based on these findings, we propose several guidelines for meta-learning new RL algorithms which will help ensure that future learned algorithms are as performant as possible.
title How Should We Meta-Learn Reinforcement Learning Algorithms?
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2507.17668