Goal-Conditioned Agents that Learn Everything All at Once

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Matthews, Michael, Jackson, Matthew, Beukman, Michael, Foster, Thomas, Letcher, Alistair, Fujimoto, Scott, Colas, Cédric, Foerster, Jakob
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913155709403136
author Matthews, Michael
Jackson, Matthew
Beukman, Michael
Foster, Thomas
Letcher, Alistair
Fujimoto, Scott
Colas, Cédric
Foerster, Jakob
author_facet Matthews, Michael
Jackson, Matthew
Beukman, Michael
Foster, Thomas
Letcher, Alistair
Fujimoto, Scott
Colas, Cédric
Foerster, Jakob
contents A goal-conditioned reinforcement learning agent exploring an environment will see a wealth of information throughout a trajectory, most of which is discarded when only performing on-policy updates with respect to the commanded goal. All-goals learning, where each transition is used for learning off-policy with respect to every goal, allows agents to extract maximal information, however it is usually computationally infeasible when done via naive relabelling. This can be overcome by jointly outputting values and actions for every goal at once, allowing for efficient, parallel all-goals updates with a single pass through the network, in a process we call Learning Everything all at Once (LEO). We show that this approach significantly outperforms other methods on goal-conditioned Craftax and is competitive with existing baselines on continuous control environments, while achieving a >250x speed-up compared to all-goals relabelling. We then go on to show that this approach can be made even more powerful by using LEO as a teacher network, rather than a direct actor. We hope that, by unlocking all-goals learning at scale, LEO can serve as a useful tool for RL practitioners in complex environments. We open source our code.
format Preprint
id arxiv_https___arxiv_org_abs_2605_23551
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Goal-Conditioned Agents that Learn Everything All at Once
Matthews, Michael
Jackson, Matthew
Beukman, Michael
Foster, Thomas
Letcher, Alistair
Fujimoto, Scott
Colas, Cédric
Foerster, Jakob
Machine Learning
Artificial Intelligence
A goal-conditioned reinforcement learning agent exploring an environment will see a wealth of information throughout a trajectory, most of which is discarded when only performing on-policy updates with respect to the commanded goal. All-goals learning, where each transition is used for learning off-policy with respect to every goal, allows agents to extract maximal information, however it is usually computationally infeasible when done via naive relabelling. This can be overcome by jointly outputting values and actions for every goal at once, allowing for efficient, parallel all-goals updates with a single pass through the network, in a process we call Learning Everything all at Once (LEO). We show that this approach significantly outperforms other methods on goal-conditioned Craftax and is competitive with existing baselines on continuous control environments, while achieving a >250x speed-up compared to all-goals relabelling. We then go on to show that this approach can be made even more powerful by using LEO as a teacher network, rather than a direct actor. We hope that, by unlocking all-goals learning at scale, LEO can serve as a useful tool for RL practitioners in complex environments. We open source our code.
title Goal-Conditioned Agents that Learn Everything All at Once
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.23551