Horizon Reduction Makes RL Scalable

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Park, Seohong, Frans, Kevin, Mann, Deepinder, Eysenbach, Benjamin, Kumar, Aviral, Levine, Sergey
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908605193650176
author Park, Seohong
Frans, Kevin
Mann, Deepinder
Eysenbach, Benjamin
Kumar, Aviral
Levine, Sergey
author_facet Park, Seohong
Frans, Kevin
Mann, Deepinder
Eysenbach, Benjamin
Kumar, Aviral
Levine, Sergey
contents In this work, we study the scalability of offline reinforcement learning (RL) algorithms. In principle, a truly scalable offline RL algorithm should be able to solve any given problem, regardless of its complexity, given sufficient data, compute, and model capacity. We investigate if and how current offline RL algorithms match up to this promise on diverse, challenging, previously unsolved tasks, using datasets up to 1000x larger than typical offline RL datasets. We observe that despite scaling up data, many existing offline RL algorithms exhibit poor scaling behavior, saturating well below the maximum performance. We hypothesize that the horizon is the main cause behind the poor scaling of offline RL. We empirically verify this hypothesis through several analysis experiments, showing that long horizons indeed present a fundamental barrier to scaling up offline RL. We then show that various horizon reduction techniques substantially enhance scalability on challenging tasks. Based on our insights, we also introduce a minimal yet scalable method named SHARSA that effectively reduces the horizon. SHARSA achieves the best asymptotic performance and scaling behavior among our evaluation methods, showing that explicitly reducing the horizon unlocks the scalability of offline RL. Code: https://github.com/seohongpark/horizon-reduction
format Preprint
id arxiv_https___arxiv_org_abs_2506_04168
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Horizon Reduction Makes RL Scalable
Park, Seohong
Frans, Kevin
Mann, Deepinder
Eysenbach, Benjamin
Kumar, Aviral
Levine, Sergey
Machine Learning
Artificial Intelligence
In this work, we study the scalability of offline reinforcement learning (RL) algorithms. In principle, a truly scalable offline RL algorithm should be able to solve any given problem, regardless of its complexity, given sufficient data, compute, and model capacity. We investigate if and how current offline RL algorithms match up to this promise on diverse, challenging, previously unsolved tasks, using datasets up to 1000x larger than typical offline RL datasets. We observe that despite scaling up data, many existing offline RL algorithms exhibit poor scaling behavior, saturating well below the maximum performance. We hypothesize that the horizon is the main cause behind the poor scaling of offline RL. We empirically verify this hypothesis through several analysis experiments, showing that long horizons indeed present a fundamental barrier to scaling up offline RL. We then show that various horizon reduction techniques substantially enhance scalability on challenging tasks. Based on our insights, we also introduce a minimal yet scalable method named SHARSA that effectively reduces the horizon. SHARSA achieves the best asymptotic performance and scaling behavior among our evaluation methods, showing that explicitly reducing the horizon unlocks the scalability of offline RL. Code: https://github.com/seohongpark/horizon-reduction
title Horizon Reduction Makes RL Scalable
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.04168