A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Svedas, Jonas, Watson, Hannah, Laubeuf, Nathan, Moolchandani, Diksha, Nada, Abubakr, Singh, Arjun, Biswas, Dwaipayan, Myers, James, Bhattacharjee, Debjyoti
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909645825638400
author Svedas, Jonas
Watson, Hannah
Laubeuf, Nathan
Moolchandani, Diksha
Nada, Abubakr
Singh, Arjun
Biswas, Dwaipayan
Myers, James
Bhattacharjee, Debjyoti
author_facet Svedas, Jonas
Watson, Hannah
Laubeuf, Nathan
Moolchandani, Diksha
Nada, Abubakr
Singh, Arjun
Biswas, Dwaipayan
Myers, James
Bhattacharjee, Debjyoti
contents Distributed deep neural networks (DNNs) have become a cornerstone for scaling machine learning to meet the demands of increasingly complex applications. However, the rapid growth in model complexity far outpaces CMOS technology scaling, making sustainable and efficient system design a critical challenge. Addressing this requires coordinated co-design across software, hardware, and technology layers. Due to the prohibitive cost and complexity of deploying full-scale training systems, simulators play a pivotal role in enabling this design exploration. This survey reviews the landscape of distributed DNN training simulators, focusing on three major dimensions: workload representation, simulation infrastructure, and models for total cost of ownership (TCO) including carbon emissions. It covers how workloads are abstracted and used in simulation, outlines common workload representation methods, and includes comprehensive comparison tables covering both simulation frameworks and TCO/emissions models, detailing their capabilities, assumptions, and areas of focus. In addition to synthesizing existing tools, the survey highlights emerging trends, common limitations, and open research challenges across the stack. By providing a structured overview, this work supports informed decision-making in the design and evaluation of distributed training systems.
format Preprint
id arxiv_https___arxiv_org_abs_2506_09275
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO
Svedas, Jonas
Watson, Hannah
Laubeuf, Nathan
Moolchandani, Diksha
Nada, Abubakr
Singh, Arjun
Biswas, Dwaipayan
Myers, James
Bhattacharjee, Debjyoti
Distributed, Parallel, and Cluster Computing
Distributed deep neural networks (DNNs) have become a cornerstone for scaling machine learning to meet the demands of increasingly complex applications. However, the rapid growth in model complexity far outpaces CMOS technology scaling, making sustainable and efficient system design a critical challenge. Addressing this requires coordinated co-design across software, hardware, and technology layers. Due to the prohibitive cost and complexity of deploying full-scale training systems, simulators play a pivotal role in enabling this design exploration. This survey reviews the landscape of distributed DNN training simulators, focusing on three major dimensions: workload representation, simulation infrastructure, and models for total cost of ownership (TCO) including carbon emissions. It covers how workloads are abstracted and used in simulation, outlines common workload representation methods, and includes comprehensive comparison tables covering both simulation frameworks and TCO/emissions models, detailing their capabilities, assumptions, and areas of focus. In addition to synthesizing existing tools, the survey highlights emerging trends, common limitations, and open research challenges across the stack. By providing a structured overview, this work supports informed decision-making in the design and evaluation of distributed training systems.
title A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2506.09275