Hybrid Approach to Parallel Stochastic Gradient Descent

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Vora, Aakash Sudhirbhai, Joshi, Dhrumil Chetankumar, Patel, Aksh Kantibhai
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909234245926912
author Vora, Aakash Sudhirbhai
Joshi, Dhrumil Chetankumar
Patel, Aksh Kantibhai
author_facet Vora, Aakash Sudhirbhai
Joshi, Dhrumil Chetankumar
Patel, Aksh Kantibhai
contents Stochastic Gradient Descent is used for large datasets to train models to reduce the training time. On top of that data parallelism is widely used as a method to efficiently train neural networks using multiple worker nodes in parallel. Synchronous and asynchronous approach to data parallelism is used by most systems to train the model in parallel. However, both of them have their drawbacks. We propose a third approach to data parallelism which is a hybrid between synchronous and asynchronous approaches, using both approaches to train the neural network. When the threshold function is selected appropriately to gradually shift all parameter aggregation from asynchronous to synchronous, we show that in a given time period our hybrid approach outperforms both asynchronous and synchronous approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2407_00101
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Hybrid Approach to Parallel Stochastic Gradient Descent
Vora, Aakash Sudhirbhai
Joshi, Dhrumil Chetankumar
Patel, Aksh Kantibhai
Machine Learning
Artificial Intelligence
Computational Complexity
Distributed, Parallel, and Cluster Computing
Neural and Evolutionary Computing
Stochastic Gradient Descent is used for large datasets to train models to reduce the training time. On top of that data parallelism is widely used as a method to efficiently train neural networks using multiple worker nodes in parallel. Synchronous and asynchronous approach to data parallelism is used by most systems to train the model in parallel. However, both of them have their drawbacks. We propose a third approach to data parallelism which is a hybrid between synchronous and asynchronous approaches, using both approaches to train the neural network. When the threshold function is selected appropriately to gradually shift all parameter aggregation from asynchronous to synchronous, we show that in a given time period our hybrid approach outperforms both asynchronous and synchronous approaches.
title Hybrid Approach to Parallel Stochastic Gradient Descent
topic Machine Learning
Artificial Intelligence
Computational Complexity
Distributed, Parallel, and Cluster Computing
Neural and Evolutionary Computing
url https://arxiv.org/abs/2407.00101