A Distributed Framework for Causal Modeling of Performance Variability in GPU Traces

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lahiry, Ankur, Pokharel, Ayush, Banday, Banooqa, Ockerman, Seth, Gueroudji, Amal, Zaeed, Mohammad, Islam, Tanzima Z., Pouchard, Line
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918164689846272
author Lahiry, Ankur
Pokharel, Ayush
Banday, Banooqa
Ockerman, Seth
Gueroudji, Amal
Zaeed, Mohammad
Islam, Tanzima Z.
Pouchard, Line
author_facet Lahiry, Ankur
Pokharel, Ayush
Banday, Banooqa
Ockerman, Seth
Gueroudji, Amal
Zaeed, Mohammad
Islam, Tanzima Z.
Pouchard, Line
contents Large-scale GPU traces play a critical role in identifying performance bottlenecks within heterogeneous High-Performance Computing (HPC) architectures. However, the sheer volume and complexity of a single trace of data make performance analysis both computationally expensive and time-consuming. To address this challenge, we present an end-to-end parallel performance analysis framework designed to handle multiple large-scale GPU traces efficiently. Our proposed framework partitions and processes trace data concurrently and employs causal graph methods and parallel coordinating chart to expose performance variability and dependencies across execution flows. Experimental results demonstrate a 67% improvement in terms of scalability, highlighting the effectiveness of our pipeline for analyzing multiple traces independently.
format Preprint
id arxiv_https___arxiv_org_abs_2510_18300
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Distributed Framework for Causal Modeling of Performance Variability in GPU Traces
Lahiry, Ankur
Pokharel, Ayush
Banday, Banooqa
Ockerman, Seth
Gueroudji, Amal
Zaeed, Mohammad
Islam, Tanzima Z.
Pouchard, Line
Distributed, Parallel, and Cluster Computing
Machine Learning
Large-scale GPU traces play a critical role in identifying performance bottlenecks within heterogeneous High-Performance Computing (HPC) architectures. However, the sheer volume and complexity of a single trace of data make performance analysis both computationally expensive and time-consuming. To address this challenge, we present an end-to-end parallel performance analysis framework designed to handle multiple large-scale GPU traces efficiently. Our proposed framework partitions and processes trace data concurrently and employs causal graph methods and parallel coordinating chart to expose performance variability and dependencies across execution flows. Experimental results demonstrate a 67% improvement in terms of scalability, highlighting the effectiveness of our pipeline for analyzing multiple traces independently.
title A Distributed Framework for Causal Modeling of Performance Variability in GPU Traces
topic Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2510.18300