Bigger, Regularized, Optimistic: scaling for compute and sample-efficient continuous control

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Nauman, Michal, Ostaszewski, Mateusz, Jankowski, Krzysztof, Miłoś, Piotr, Cygan, Marek
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916503564058624
author Nauman, Michal
Ostaszewski, Mateusz
Jankowski, Krzysztof
Miłoś, Piotr
Cygan, Marek
author_facet Nauman, Michal
Ostaszewski, Mateusz
Jankowski, Krzysztof
Miłoś, Piotr
Cygan, Marek
contents Sample efficiency in Reinforcement Learning (RL) has traditionally been driven by algorithmic enhancements. In this work, we demonstrate that scaling can also lead to substantial improvements. We conduct a thorough investigation into the interplay of scaling model capacity and domain-specific RL enhancements. These empirical findings inform the design choices underlying our proposed BRO (Bigger, Regularized, Optimistic) algorithm. The key innovation behind BRO is that strong regularization allows for effective scaling of the critic networks, which, paired with optimistic exploration, leads to superior performance. BRO achieves state-of-the-art results, significantly outperforming the leading model-based and model-free algorithms across 40 complex tasks from the DeepMind Control, MetaWorld, and MyoSuite benchmarks. BRO is the first model-free algorithm to achieve near-optimal policies in the notoriously challenging Dog and Humanoid tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2405_16158
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Bigger, Regularized, Optimistic: scaling for compute and sample-efficient continuous control
Nauman, Michal
Ostaszewski, Mateusz
Jankowski, Krzysztof
Miłoś, Piotr
Cygan, Marek
Machine Learning
Sample efficiency in Reinforcement Learning (RL) has traditionally been driven by algorithmic enhancements. In this work, we demonstrate that scaling can also lead to substantial improvements. We conduct a thorough investigation into the interplay of scaling model capacity and domain-specific RL enhancements. These empirical findings inform the design choices underlying our proposed BRO (Bigger, Regularized, Optimistic) algorithm. The key innovation behind BRO is that strong regularization allows for effective scaling of the critic networks, which, paired with optimistic exploration, leads to superior performance. BRO achieves state-of-the-art results, significantly outperforming the leading model-based and model-free algorithms across 40 complex tasks from the DeepMind Control, MetaWorld, and MyoSuite benchmarks. BRO is the first model-free algorithm to achieve near-optimal policies in the notoriously challenging Dog and Humanoid tasks.
title Bigger, Regularized, Optimistic: scaling for compute and sample-efficient continuous control
topic Machine Learning
url https://arxiv.org/abs/2405.16158