Scalable and Efficient Distributed Training of Deep Learning Models via Hybrid Parallelism

Fuente: Zenodo
Guardado en:
Detalles Bibliográficos
Autor principal: Yuki Tanaka
Formato: Recurso digital
Publicado: Zenodo 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866901471269748736
author Yuki Tanaka
author_facet Yuki Tanaka
contents Deep learning models have achieved remarkable success in various domains, but their training often demands significant computational resources. This paper investigates the application of hybrid parallelism, combining data and model parallelism, to enhance the scalability and efficiency of distributed deep learning training. We present a performance analysis of different hybrid parallelism strategies on a high-performance computing cluster, highlighting the trade-offs between communication overhead and computational workload distribution. Our results demonstrate that a carefully tuned hybrid parallelism approach can significantly reduce training time and improve resource utilization compared to pure data or model parallelism.
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_18945485
institution Zenodo
language
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle Scalable and Efficient Distributed Training of Deep Learning Models via Hybrid Parallelism
Yuki Tanaka
machine learning
deep learning
artificial intelligence
Deep learning models have achieved remarkable success in various domains, but their training often demands significant computational resources. This paper investigates the application of hybrid parallelism, combining data and model parallelism, to enhance the scalability and efficiency of distributed deep learning training. We present a performance analysis of different hybrid parallelism strategies on a high-performance computing cluster, highlighting the trade-offs between communication overhead and computational workload distribution. Our results demonstrate that a carefully tuned hybrid parallelism approach can significantly reduce training time and improve resource utilization compared to pure data or model parallelism.
title Scalable and Efficient Distributed Training of Deep Learning Models via Hybrid Parallelism
topic machine learning
deep learning
artificial intelligence
url https://doi.org/10.5281/zenodo.18945485