Extracting Practical, Actionable Energy Insights from Supercomputer Telemetry and Logs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cornelius, Melanie, Cross, Greg, Shilpika, Shilpika, Dearing, Matthew T., Lan, Zhiling
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908372788314112
author Cornelius, Melanie
Cross, Greg
Shilpika, Shilpika
Dearing, Matthew T.
Lan, Zhiling
author_facet Cornelius, Melanie
Cross, Greg
Shilpika, Shilpika
Dearing, Matthew T.
Lan, Zhiling
contents As supercomputers grow in size and complexity, power efficiency has become a critical challenge, particularly in understanding GPU power consumption within modern HPC workloads. This work addresses this challenge by presenting a data co-analysis approach using system data collected from the Polaris supercomputer at Argonne National Laboratory. We focus on GPU utilization and power demands, navigating the complexities of large-scale, heterogeneous datasets. Our approach, which incorporates data preprocessing, post-processing, and statistical methods, condenses the data volume by 94% while preserving essential insights. Through this analysis, we uncover key opportunities for power optimization, such as reducing high idle power costs, applying power strategies at the job-level, and aligning GPU power allocation with workload demands. Our findings provide actionable insights for energy-efficient computing and offer a practical, reproducible approach for applying existing research to optimize system performance.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14796
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Extracting Practical, Actionable Energy Insights from Supercomputer Telemetry and Logs
Cornelius, Melanie
Cross, Greg
Shilpika, Shilpika
Dearing, Matthew T.
Lan, Zhiling
Distributed, Parallel, and Cluster Computing
Performance
As supercomputers grow in size and complexity, power efficiency has become a critical challenge, particularly in understanding GPU power consumption within modern HPC workloads. This work addresses this challenge by presenting a data co-analysis approach using system data collected from the Polaris supercomputer at Argonne National Laboratory. We focus on GPU utilization and power demands, navigating the complexities of large-scale, heterogeneous datasets. Our approach, which incorporates data preprocessing, post-processing, and statistical methods, condenses the data volume by 94% while preserving essential insights. Through this analysis, we uncover key opportunities for power optimization, such as reducing high idle power costs, applying power strategies at the job-level, and aligning GPU power allocation with workload demands. Our findings provide actionable insights for energy-efficient computing and offer a practical, reproducible approach for applying existing research to optimize system performance.
title Extracting Practical, Actionable Energy Insights from Supercomputer Telemetry and Logs
topic Distributed, Parallel, and Cluster Computing
Performance
url https://arxiv.org/abs/2505.14796