Hybrid Approach to Parallel Stochastic Gradient Descent
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866909234245926912 |
|---|---|
| author | Vora, Aakash Sudhirbhai Joshi, Dhrumil Chetankumar Patel, Aksh Kantibhai |
| author_facet | Vora, Aakash Sudhirbhai Joshi, Dhrumil Chetankumar Patel, Aksh Kantibhai |
| contents | Stochastic Gradient Descent is used for large datasets to train models to reduce the training time. On top of that data parallelism is widely used as a method to efficiently train neural networks using multiple worker nodes in parallel. Synchronous and asynchronous approach to data parallelism is used by most systems to train the model in parallel. However, both of them have their drawbacks. We propose a third approach to data parallelism which is a hybrid between synchronous and asynchronous approaches, using both approaches to train the neural network. When the threshold function is selected appropriately to gradually shift all parameter aggregation from asynchronous to synchronous, we show that in a given time period our hybrid approach outperforms both asynchronous and synchronous approaches. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_00101 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Hybrid Approach to Parallel Stochastic Gradient Descent Vora, Aakash Sudhirbhai Joshi, Dhrumil Chetankumar Patel, Aksh Kantibhai Machine Learning Artificial Intelligence Computational Complexity Distributed, Parallel, and Cluster Computing Neural and Evolutionary Computing Stochastic Gradient Descent is used for large datasets to train models to reduce the training time. On top of that data parallelism is widely used as a method to efficiently train neural networks using multiple worker nodes in parallel. Synchronous and asynchronous approach to data parallelism is used by most systems to train the model in parallel. However, both of them have their drawbacks. We propose a third approach to data parallelism which is a hybrid between synchronous and asynchronous approaches, using both approaches to train the neural network. When the threshold function is selected appropriately to gradually shift all parameter aggregation from asynchronous to synchronous, we show that in a given time period our hybrid approach outperforms both asynchronous and synchronous approaches. |
| title | Hybrid Approach to Parallel Stochastic Gradient Descent |
| topic | Machine Learning Artificial Intelligence Computational Complexity Distributed, Parallel, and Cluster Computing Neural and Evolutionary Computing |
| url | https://arxiv.org/abs/2407.00101 |