A Comprehensive Study of Bugs in Modern Distributed Deep Learning Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Xiaoxue, Zhan, Wanwei, Chen, Jiale, Li, Yishu, Keung, Jacky, Sarro, Federica
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911335214743552
author Ma, Xiaoxue
Zhan, Wanwei
Chen, Jiale
Li, Yishu
Keung, Jacky
Sarro, Federica
author_facet Ma, Xiaoxue
Zhan, Wanwei
Chen, Jiale
Li, Yishu
Keung, Jacky
Sarro, Federica
contents In today's data-driven era, deep learning is vital for processing massive datasets, yet single-device training is constrained by computational and memory limits. Distributed deep learning overcomes these challenges by leveraging multiple GPUs or machines in parallel. While general-purpose frameworks (e.g., TensorFlow and PyTorch) provide distributed capabilities, these are often add-on features that demand significant manual effort for advanced parallelism, underscoring the need for specialized frameworks. This study conducts the first large-scale empirical analysis of practitioner challenges in dedicated distributed frameworks. We examine 849 real-world issues from DeepSpeed, Megatron-LM, and Colossal-AI and construct a taxonomy of 34 bug symptoms, 28 root causes, and 6 fix patterns. Crucially, we establish explicit mappings between symptoms, causes, and fixes across distributed training stages, enabling a systematic understanding of how issues emerge and are resolved. Our results show that 45.1\% of bug symptoms are unique to distributed frameworks, with setup failures, memory issues, and performance anomalies being the most prevalent. Moreover, 95\% of issues in the communication setup stage occur exclusively in distributed contexts. We also find over 60\% of cases can be resolved through version and dependency management, and distributed feature, API, and communication tuning. Based on these findings, we provide actionable implications.
format Preprint
id arxiv_https___arxiv_org_abs_2512_20345
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Comprehensive Study of Bugs in Modern Distributed Deep Learning Systems
Ma, Xiaoxue
Zhan, Wanwei
Chen, Jiale
Li, Yishu
Keung, Jacky
Sarro, Federica
Software Engineering
In today's data-driven era, deep learning is vital for processing massive datasets, yet single-device training is constrained by computational and memory limits. Distributed deep learning overcomes these challenges by leveraging multiple GPUs or machines in parallel. While general-purpose frameworks (e.g., TensorFlow and PyTorch) provide distributed capabilities, these are often add-on features that demand significant manual effort for advanced parallelism, underscoring the need for specialized frameworks. This study conducts the first large-scale empirical analysis of practitioner challenges in dedicated distributed frameworks. We examine 849 real-world issues from DeepSpeed, Megatron-LM, and Colossal-AI and construct a taxonomy of 34 bug symptoms, 28 root causes, and 6 fix patterns. Crucially, we establish explicit mappings between symptoms, causes, and fixes across distributed training stages, enabling a systematic understanding of how issues emerge and are resolved. Our results show that 45.1\% of bug symptoms are unique to distributed frameworks, with setup failures, memory issues, and performance anomalies being the most prevalent. Moreover, 95\% of issues in the communication setup stage occur exclusively in distributed contexts. We also find over 60\% of cases can be resolved through version and dependency management, and distributed feature, API, and communication tuning. Based on these findings, we provide actionable implications.
title A Comprehensive Study of Bugs in Modern Distributed Deep Learning Systems
topic Software Engineering
url https://arxiv.org/abs/2512.20345