A Distributed Framework for Causal Modeling of Performance Variability in GPU Traces
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866918164689846272 |
|---|---|
| author | Lahiry, Ankur Pokharel, Ayush Banday, Banooqa Ockerman, Seth Gueroudji, Amal Zaeed, Mohammad Islam, Tanzima Z. Pouchard, Line |
| author_facet | Lahiry, Ankur Pokharel, Ayush Banday, Banooqa Ockerman, Seth Gueroudji, Amal Zaeed, Mohammad Islam, Tanzima Z. Pouchard, Line |
| contents | Large-scale GPU traces play a critical role in identifying performance bottlenecks within heterogeneous High-Performance Computing (HPC) architectures. However, the sheer volume and complexity of a single trace of data make performance analysis both computationally expensive and time-consuming. To address this challenge, we present an end-to-end parallel performance analysis framework designed to handle multiple large-scale GPU traces efficiently. Our proposed framework partitions and processes trace data concurrently and employs causal graph methods and parallel coordinating chart to expose performance variability and dependencies across execution flows. Experimental results demonstrate a 67% improvement in terms of scalability, highlighting the effectiveness of our pipeline for analyzing multiple traces independently. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_18300 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | A Distributed Framework for Causal Modeling of Performance Variability in GPU Traces Lahiry, Ankur Pokharel, Ayush Banday, Banooqa Ockerman, Seth Gueroudji, Amal Zaeed, Mohammad Islam, Tanzima Z. Pouchard, Line Distributed, Parallel, and Cluster Computing Machine Learning Large-scale GPU traces play a critical role in identifying performance bottlenecks within heterogeneous High-Performance Computing (HPC) architectures. However, the sheer volume and complexity of a single trace of data make performance analysis both computationally expensive and time-consuming. To address this challenge, we present an end-to-end parallel performance analysis framework designed to handle multiple large-scale GPU traces efficiently. Our proposed framework partitions and processes trace data concurrently and employs causal graph methods and parallel coordinating chart to expose performance variability and dependencies across execution flows. Experimental results demonstrate a 67% improvement in terms of scalability, highlighting the effectiveness of our pipeline for analyzing multiple traces independently. |
| title | A Distributed Framework for Causal Modeling of Performance Variability in GPU Traces |
| topic | Distributed, Parallel, and Cluster Computing Machine Learning |
| url | https://arxiv.org/abs/2510.18300 |