AllReduce Scheduling with Hierarchical Deep Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wei, Yufan, Liu, Mickel, Wu, Wenfei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912296465334272
author Wei, Yufan
Liu, Mickel
Wu, Wenfei
author_facet Wei, Yufan
Liu, Mickel
Wu, Wenfei
contents AllReduce is a technique in distributed computing which saw use in many critical applications of deep learning. Existing methods of AllReduce scheduling oftentimes lack flexibility due to being topology-specific or relying on extensive handcrafted designs that require domain-specific knowledge. In this work, we aim to alleviate this inflexibility by proposing a deep-reinforcement-learning (DRL)-based pipeline that can generate AllReduce scheduling for various network topologies without topology-specific design features. The flow scheduling module of this pipeline consists of two hierarchically-structured DRL policies that work cooperatively to find optimal scheduling. We showcase the performance of our method compared to the baseline methods on three topologies: BCube, DCell, and Jellyfish. Finally, we contributed a Python-based simulation environment simulating AllReduce scheduling on these network topologies.
format Preprint
id arxiv_https___arxiv_org_abs_2503_21013
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AllReduce Scheduling with Hierarchical Deep Reinforcement Learning
Wei, Yufan
Liu, Mickel
Wu, Wenfei
Networking and Internet Architecture
Distributed, Parallel, and Cluster Computing
AllReduce is a technique in distributed computing which saw use in many critical applications of deep learning. Existing methods of AllReduce scheduling oftentimes lack flexibility due to being topology-specific or relying on extensive handcrafted designs that require domain-specific knowledge. In this work, we aim to alleviate this inflexibility by proposing a deep-reinforcement-learning (DRL)-based pipeline that can generate AllReduce scheduling for various network topologies without topology-specific design features. The flow scheduling module of this pipeline consists of two hierarchically-structured DRL policies that work cooperatively to find optimal scheduling. We showcase the performance of our method compared to the baseline methods on three topologies: BCube, DCell, and Jellyfish. Finally, we contributed a Python-based simulation environment simulating AllReduce scheduling on these network topologies.
title AllReduce Scheduling with Hierarchical Deep Reinforcement Learning
topic Networking and Internet Architecture
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2503.21013