Efficient Fine-Grained GPU Performance Modeling for Distributed Deep Learning of LLM

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Biyao, Zheng, Mingkai, Ganguly, Debargha, Zhang, Xuecen, Singh, Vikash, Chaudhary, Vipin, Zhang, Zhao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909810404884480
author Zhang, Biyao
Zheng, Mingkai
Ganguly, Debargha
Zhang, Xuecen
Singh, Vikash
Chaudhary, Vipin
Zhang, Zhao
author_facet Zhang, Biyao
Zheng, Mingkai
Ganguly, Debargha
Zhang, Xuecen
Singh, Vikash
Chaudhary, Vipin
Zhang, Zhao
contents Training Large Language Models(LLMs) is one of the most compute-intensive tasks in high-performance computing. Predicting end-to-end training time for multi-billion parameter models distributed across hundreds of GPUs remains challenging due to complex interactions between transformer components, parallelism strategies(data, model, pipeline, tensor), and multi-tier communication. Learned models require costly sampling, while analytical models often struggle with real-world network and hardware complexities. We address this by decomposing LLMs into core computational primitives and modeling them with: (1) operator-level decomposition for fine-grained analysis; (2) lightweight sampling based hardware-aware prediction models for key operations; (3) an end-to-end prediction system integrating these components across complex parallelization strategies. Crucially, our methodology has been validated on two large-scale HPC systems. Our framework achieves low average prediction errors-4.98\% on Perlmutter(A100) and 9.38\% on Vista(GH200)-for models up to 20B parameters across 128 GPUs. Importantly, it runs entirely on CPUs, enabling rapid iteration over hardware configurations and training strategies without costly on-cluster experimentation.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22832
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient Fine-Grained GPU Performance Modeling for Distributed Deep Learning of LLM
Zhang, Biyao
Zheng, Mingkai
Ganguly, Debargha
Zhang, Xuecen
Singh, Vikash
Chaudhary, Vipin
Zhang, Zhao
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
Training Large Language Models(LLMs) is one of the most compute-intensive tasks in high-performance computing. Predicting end-to-end training time for multi-billion parameter models distributed across hundreds of GPUs remains challenging due to complex interactions between transformer components, parallelism strategies(data, model, pipeline, tensor), and multi-tier communication. Learned models require costly sampling, while analytical models often struggle with real-world network and hardware complexities. We address this by decomposing LLMs into core computational primitives and modeling them with: (1) operator-level decomposition for fine-grained analysis; (2) lightweight sampling based hardware-aware prediction models for key operations; (3) an end-to-end prediction system integrating these components across complex parallelization strategies. Crucially, our methodology has been validated on two large-scale HPC systems. Our framework achieves low average prediction errors-4.98\% on Perlmutter(A100) and 9.38\% on Vista(GH200)-for models up to 20B parameters across 128 GPUs. Importantly, it runs entirely on CPUs, enabling rapid iteration over hardware configurations and training strategies without costly on-cluster experimentation.
title Efficient Fine-Grained GPU Performance Modeling for Distributed Deep Learning of LLM
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.22832