SRL: Scaling Distributed Reinforcement Learning to Over Ten Thousand Cores

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mei, Zhiyu, Fu, Wei, Gao, Jiaxuan, Wang, Guangju, Zhang, Huanchen, Wu, Yi
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910496244891648
author Mei, Zhiyu
Fu, Wei
Gao, Jiaxuan
Wang, Guangju
Zhang, Huanchen
Wu, Yi
author_facet Mei, Zhiyu
Fu, Wei
Gao, Jiaxuan
Wang, Guangju
Zhang, Huanchen
Wu, Yi
contents The ever-growing complexity of reinforcement learning (RL) tasks demands a distributed system to efficiently generate and process a massive amount of data. However, existing open-source libraries suffer from various limitations, which impede their practical use in challenging scenarios where large-scale training is necessary. In this paper, we present a novel abstraction on the dataflows of RL training, which unifies diverse RL training applications into a general framework. Following this abstraction, we develop a scalable, efficient, and extensible distributed RL system called ReaLlyScalableRL, which allows efficient and massively parallelized training and easy development of customized algorithms. Our evaluation shows that SRL outperforms existing academic libraries, reaching at most 21x higher training throughput in a distributed setting. On learning performance, beyond performing and scaling well on common RL benchmarks with different RL algorithms, SRL can reproduce the same solution in the challenging hide-and-seek environment as reported by OpenAI with up to 5x speedup in wall-clock time. Notably, SRL is the first in the academic community to perform RL experiments at a large scale with over 15k CPU cores. SRL source code is available at: https://github.com/openpsi-project/srl .
format Preprint
id arxiv_https___arxiv_org_abs_2306_16688
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle SRL: Scaling Distributed Reinforcement Learning to Over Ten Thousand Cores
Mei, Zhiyu
Fu, Wei
Gao, Jiaxuan
Wang, Guangju
Zhang, Huanchen
Wu, Yi
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
The ever-growing complexity of reinforcement learning (RL) tasks demands a distributed system to efficiently generate and process a massive amount of data. However, existing open-source libraries suffer from various limitations, which impede their practical use in challenging scenarios where large-scale training is necessary. In this paper, we present a novel abstraction on the dataflows of RL training, which unifies diverse RL training applications into a general framework. Following this abstraction, we develop a scalable, efficient, and extensible distributed RL system called ReaLlyScalableRL, which allows efficient and massively parallelized training and easy development of customized algorithms. Our evaluation shows that SRL outperforms existing academic libraries, reaching at most 21x higher training throughput in a distributed setting. On learning performance, beyond performing and scaling well on common RL benchmarks with different RL algorithms, SRL can reproduce the same solution in the challenging hide-and-seek environment as reported by OpenAI with up to 5x speedup in wall-clock time. Notably, SRL is the first in the academic community to perform RL experiments at a large scale with over 15k CPU cores. SRL source code is available at: https://github.com/openpsi-project/srl .
title SRL: Scaling Distributed Reinforcement Learning to Over Ten Thousand Cores
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2306.16688