Fine-Grained Power and Energy Attribution on AMD GPU/APU-Based Exascale Nodes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: McDaniel, Adam, Jantz, Michael, Sharma, Ashesh, Abbott, Steve, Martin, Steven, Khandekar, Shreyas, Neth, Brandon, Alvarez, Bruno Villasenor, Kashi, Aditya, Elwasif, Wael, Hernandez, Oscar
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915929363841024
author McDaniel, Adam
Jantz, Michael
Sharma, Ashesh
Abbott, Steve
Martin, Steven
Khandekar, Shreyas
Neth, Brandon
Alvarez, Bruno Villasenor
Kashi, Aditya
Elwasif, Wael
Hernandez, Oscar
author_facet McDaniel, Adam
Jantz, Michael
Sharma, Ashesh
Abbott, Steve
Martin, Steven
Khandekar, Shreyas
Neth, Brandon
Alvarez, Bruno Villasenor
Kashi, Aditya
Elwasif, Wael
Hernandez, Oscar
contents Modern exascale GPU- and APU-based systems provide multiple power and energy sensors, but differences in scope, update rate, timing, and filtering complicate the attribution of short-lived accelerator activity. This paper presents a methodology to characterize and correct these effects on Cray EX systems with AMD Instinct MI250X GPUs (Frontier) and MI300A APUs (Portage). Using controlled square-wave workloads, we quantify update intervals, delay, aliasing, and variability across up to 512 GPUs and 480 APUs with on-chip (rocm-smi/amd-smi) and off-chip Cray Power Management sensors. We reconstruct power from cumulative energy counters to achieve faster response times, validate it against on-chip, off-chip, and node-level sensors, and integrate the resulting streams into a Score-P/PAPI-based tool for time-aligned, phase-level attribution. Applied to rocHPL, rocHPL-MxP, and HPG-MxP, the method separates energy savings due to reduced runtime from changes in power. Mixed precision reduces node energy on Frontier by 79% for rocHPL-MxP and 31% for HPG-MxP, with similar trends on Portage. These results provide portable guidance for sensor validation and power-aware optimization on current and future exascale systems.
format Preprint
id arxiv_https___arxiv_org_abs_2604_06056
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Fine-Grained Power and Energy Attribution on AMD GPU/APU-Based Exascale Nodes
McDaniel, Adam
Jantz, Michael
Sharma, Ashesh
Abbott, Steve
Martin, Steven
Khandekar, Shreyas
Neth, Brandon
Alvarez, Bruno Villasenor
Kashi, Aditya
Elwasif, Wael
Hernandez, Oscar
Distributed, Parallel, and Cluster Computing
Hardware Architecture
Modern exascale GPU- and APU-based systems provide multiple power and energy sensors, but differences in scope, update rate, timing, and filtering complicate the attribution of short-lived accelerator activity. This paper presents a methodology to characterize and correct these effects on Cray EX systems with AMD Instinct MI250X GPUs (Frontier) and MI300A APUs (Portage). Using controlled square-wave workloads, we quantify update intervals, delay, aliasing, and variability across up to 512 GPUs and 480 APUs with on-chip (rocm-smi/amd-smi) and off-chip Cray Power Management sensors. We reconstruct power from cumulative energy counters to achieve faster response times, validate it against on-chip, off-chip, and node-level sensors, and integrate the resulting streams into a Score-P/PAPI-based tool for time-aligned, phase-level attribution. Applied to rocHPL, rocHPL-MxP, and HPG-MxP, the method separates energy savings due to reduced runtime from changes in power. Mixed precision reduces node energy on Frontier by 79% for rocHPL-MxP and 31% for HPG-MxP, with similar trends on Portage. These results provide portable guidance for sensor validation and power-aware optimization on current and future exascale systems.
title Fine-Grained Power and Energy Attribution on AMD GPU/APU-Based Exascale Nodes
topic Distributed, Parallel, and Cluster Computing
Hardware Architecture
url https://arxiv.org/abs/2604.06056