Accelerate Model Parallel Training by Using Efficient Graph Traversal Order in Device Placement

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Tianze, Payberah, Amir H., Hagos, Desta Haileselassie, Vlassov, Vladimir
Format: Preprint
Veröffentlicht: 2022
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913564051111936
author Wang, Tianze
Payberah, Amir H.
Hagos, Desta Haileselassie
Vlassov, Vladimir
author_facet Wang, Tianze
Payberah, Amir H.
Hagos, Desta Haileselassie
Vlassov, Vladimir
contents Modern neural networks require long training to reach decent performance on massive datasets. One common approach to speed up training is model parallelization, where large neural networks are split across multiple devices. However, different device placements of the same neural network lead to different training times. Most of the existing device placement solutions treat the problem as sequential decision-making by traversing neural network graphs and assigning their neurons to different devices. This work studies the impact of graph traversal order on device placement. In particular, we empirically study how different graph traversal order leads to different device placement, which in turn affects the training execution time. Our experiment results show that the best graph traversal order depends on the type of neural networks and their computation graphs features. In this work, we also provide recommendations on choosing graph traversal order in device placement for various neural network families to improve the training time in model parallelization.
format Preprint
id arxiv_https___arxiv_org_abs_2201_09676
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Accelerate Model Parallel Training by Using Efficient Graph Traversal Order in Device Placement
Wang, Tianze
Payberah, Amir H.
Hagos, Desta Haileselassie
Vlassov, Vladimir
Machine Learning
Artificial Intelligence
Modern neural networks require long training to reach decent performance on massive datasets. One common approach to speed up training is model parallelization, where large neural networks are split across multiple devices. However, different device placements of the same neural network lead to different training times. Most of the existing device placement solutions treat the problem as sequential decision-making by traversing neural network graphs and assigning their neurons to different devices. This work studies the impact of graph traversal order on device placement. In particular, we empirically study how different graph traversal order leads to different device placement, which in turn affects the training execution time. Our experiment results show that the best graph traversal order depends on the type of neural networks and their computation graphs features. In this work, we also provide recommendations on choosing graph traversal order in device placement for various neural network families to improve the training time in model parallelization.
title Accelerate Model Parallel Training by Using Efficient Graph Traversal Order in Device Placement
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2201.09676