Towards a Standardized Representation for Deep Learning Collective Algorithms

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yoo, Jinsun, Won, William, Cowan, Meghan, Jiang, Nan, Klenk, Benjamin, Sridharan, Srinivas, Krishna, Tushar
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911218631966720
author Yoo, Jinsun
Won, William
Cowan, Meghan
Jiang, Nan
Klenk, Benjamin
Sridharan, Srinivas
Krishna, Tushar
author_facet Yoo, Jinsun
Won, William
Cowan, Meghan
Jiang, Nan
Klenk, Benjamin
Sridharan, Srinivas
Krishna, Tushar
contents The explosion of machine learning model size has led to its execution on distributed clusters at a very large scale. Many works have tried to optimize the process of producing collective algorithms and running collective communications, which act as a bottleneck to distributed machine learning. However, different works use their own collective algorithm representation, pushing away from co-optimizing collective communication and the rest of the workload. The lack of a standardized collective algorithm representation has also hindered interoperability between collective algorithm producers and consumers. Additionally, tool-specific conversions and modifications have to be made for each pair of tools producing and consuming collective algorithms which adds to engineering efforts. In this position paper, we propose a standardized workflow leveraging a common collective algorithm representation. Upstream producers and downstream consumers converge to a common representation format based on Chakra Execution Trace, a commonly used graph based representation of distributed machine learning workloads. Such a common representation enables us to view collective communications at the same level as workload operations and decouple producer and consumer tools, enhance interoperability, and relieve the user from the burden of having to focus on downstream implementations. We provide a proof-of-concept of this standardized workflow by simulating collective algorithms generated by the MSCCLang domain-specific language through the ASTRA-sim distributed machine learning simulator using various network configurations.
format Preprint
id arxiv_https___arxiv_org_abs_2408_11008
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards a Standardized Representation for Deep Learning Collective Algorithms
Yoo, Jinsun
Won, William
Cowan, Meghan
Jiang, Nan
Klenk, Benjamin
Sridharan, Srinivas
Krishna, Tushar
Distributed, Parallel, and Cluster Computing
The explosion of machine learning model size has led to its execution on distributed clusters at a very large scale. Many works have tried to optimize the process of producing collective algorithms and running collective communications, which act as a bottleneck to distributed machine learning. However, different works use their own collective algorithm representation, pushing away from co-optimizing collective communication and the rest of the workload. The lack of a standardized collective algorithm representation has also hindered interoperability between collective algorithm producers and consumers. Additionally, tool-specific conversions and modifications have to be made for each pair of tools producing and consuming collective algorithms which adds to engineering efforts. In this position paper, we propose a standardized workflow leveraging a common collective algorithm representation. Upstream producers and downstream consumers converge to a common representation format based on Chakra Execution Trace, a commonly used graph based representation of distributed machine learning workloads. Such a common representation enables us to view collective communications at the same level as workload operations and decouple producer and consumer tools, enhance interoperability, and relieve the user from the burden of having to focus on downstream implementations. We provide a proof-of-concept of this standardized workflow by simulating collective algorithms generated by the MSCCLang domain-specific language through the ASTRA-sim distributed machine learning simulator using various network configurations.
title Towards a Standardized Representation for Deep Learning Collective Algorithms
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2408.11008