Scalable GPU Performance Variability Analysis framework

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lahiry, Ankur, Pokharel, Ayush, Ockerman, Seth, Gueroudji, Amal, Pouchard, Line, Islam, Tanzima Z.
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911022486388736
author Lahiry, Ankur
Pokharel, Ayush
Ockerman, Seth
Gueroudji, Amal
Pouchard, Line
Islam, Tanzima Z.
author_facet Lahiry, Ankur
Pokharel, Ayush
Ockerman, Seth
Gueroudji, Amal
Pouchard, Line
Islam, Tanzima Z.
contents Analyzing large-scale performance logs from GPU profilers often requires terabytes of memory and hours of runtime, even for basic summaries. These constraints prevent timely insight and hinder the integration of performance analytics into automated workflows. Existing analysis tools typically process data sequentially, making them ill-suited for HPC workflows with growing trace complexity and volume. We introduce a distributed data analysis framework that scales with dataset size and compute availability. Rather than treating the dataset as a single entity, our system partitions it into independently analyzable shards and processes them concurrently across MPI ranks. This design reduces per-node memory pressure, avoids central bottlenecks, and enables low-latency exploration of high-dimensional trace data. We apply the framework to end-to-end Nsight Compute traces from real HPC and AI workloads, demonstrate its ability to diagnose performance variability, and uncover the impact of memory transfer latency on GPU kernel behavior.
format Preprint
id arxiv_https___arxiv_org_abs_2506_20674
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scalable GPU Performance Variability Analysis framework
Lahiry, Ankur
Pokharel, Ayush
Ockerman, Seth
Gueroudji, Amal
Pouchard, Line
Islam, Tanzima Z.
Distributed, Parallel, and Cluster Computing
Performance
Analyzing large-scale performance logs from GPU profilers often requires terabytes of memory and hours of runtime, even for basic summaries. These constraints prevent timely insight and hinder the integration of performance analytics into automated workflows. Existing analysis tools typically process data sequentially, making them ill-suited for HPC workflows with growing trace complexity and volume. We introduce a distributed data analysis framework that scales with dataset size and compute availability. Rather than treating the dataset as a single entity, our system partitions it into independently analyzable shards and processes them concurrently across MPI ranks. This design reduces per-node memory pressure, avoids central bottlenecks, and enables low-latency exploration of high-dimensional trace data. We apply the framework to end-to-end Nsight Compute traces from real HPC and AI workloads, demonstrate its ability to diagnose performance variability, and uncover the impact of memory transfer latency on GPU kernel behavior.
title Scalable GPU Performance Variability Analysis framework
topic Distributed, Parallel, and Cluster Computing
Performance
url https://arxiv.org/abs/2506.20674