Scalable and Efficient Distributed Training of Deep Learning Models via Hybrid Parallelism
Fuente:
Zenodo
Guardado en:
| Autor principal: | |
|---|---|
| Formato: | Recurso digital |
| Publicado: |
Zenodo
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866901471269748736 |
|---|---|
| author | Yuki Tanaka |
| author_facet | Yuki Tanaka |
| contents | Deep learning models have achieved remarkable success in various domains, but their training often demands significant computational resources. This paper investigates the application of hybrid parallelism, combining data and model parallelism, to enhance the scalability and efficiency of distributed deep learning training. We present a performance analysis of different hybrid parallelism strategies on a high-performance computing cluster, highlighting the trade-offs between communication overhead and computational workload distribution. Our results demonstrate that a carefully tuned hybrid parallelism approach can significantly reduce training time and improve resource utilization compared to pure data or model parallelism. |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_18945485 |
| institution | Zenodo |
| language | |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Scalable and Efficient Distributed Training of Deep Learning Models via Hybrid Parallelism Yuki Tanaka machine learning deep learning artificial intelligence Deep learning models have achieved remarkable success in various domains, but their training often demands significant computational resources. This paper investigates the application of hybrid parallelism, combining data and model parallelism, to enhance the scalability and efficiency of distributed deep learning training. We present a performance analysis of different hybrid parallelism strategies on a high-performance computing cluster, highlighting the trade-offs between communication overhead and computational workload distribution. Our results demonstrate that a carefully tuned hybrid parallelism approach can significantly reduce training time and improve resource utilization compared to pure data or model parallelism. |
| title | Scalable and Efficient Distributed Training of Deep Learning Models via Hybrid Parallelism |
| topic | machine learning deep learning artificial intelligence |
| url | https://doi.org/10.5281/zenodo.18945485 |