Enregistré dans:
Détails bibliographiques
Auteurs principaux: Mayr, Martin, Wind, Sebastian, Schröder, Lukas, Hager, Georg, Köstler, Harald, Wellein, Gerhard
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:https://arxiv.org/abs/2603.16164
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912970754228224
author Mayr, Martin
Wind, Sebastian
Schröder, Lukas
Hager, Georg
Köstler, Harald
Wellein, Gerhard
author_facet Mayr, Martin
Wind, Sebastian
Schröder, Lukas
Hager, Georg
Köstler, Harald
Wellein, Gerhard
contents Artificial Intelligence (AI) workloads drive a rapid expansion of high-performance computing (HPC) infrastructures and increase their power and energy demands towards a critical level. AI benchmarks representing state-of-the art workloads and their understanding in the context of performance-energy trade-offs are critical to deploy efficient infrastructures and can guide energy efficiency measures, such as power capping. We introduce a benchmarking framework with popular deep learning applications from computer vision (image classification and generation) and large language models (continued pre-training and inference) implementing modern methods. Our performance analysis focuses on throughput rather than time to "completion", which is the standard metric in HPC. We analyse performance and energy efficiency under various power capping scenarios on NVIDIA H100, NVIDIA H200, and AMD MI300X GPUs. Our results reveal that no universal optimal power cap exists, as the efficiency peak varies across application types and GPU architectures. Interestingly, the two NVIDIA GPUs which mainly differ in their HBM configuration show qualitatively different performance-energy trade-offs. The developed benchmarking framework will be released as a public tool.
format Preprint
id arxiv_https___arxiv_org_abs_2603_16164
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AI Application Benchmarking: Power-Aware Performance Analysis for Vision and Language Models
Mayr, Martin
Wind, Sebastian
Schröder, Lukas
Hager, Georg
Köstler, Harald
Wellein, Gerhard
Performance
Artificial Intelligence (AI) workloads drive a rapid expansion of high-performance computing (HPC) infrastructures and increase their power and energy demands towards a critical level. AI benchmarks representing state-of-the art workloads and their understanding in the context of performance-energy trade-offs are critical to deploy efficient infrastructures and can guide energy efficiency measures, such as power capping. We introduce a benchmarking framework with popular deep learning applications from computer vision (image classification and generation) and large language models (continued pre-training and inference) implementing modern methods. Our performance analysis focuses on throughput rather than time to "completion", which is the standard metric in HPC. We analyse performance and energy efficiency under various power capping scenarios on NVIDIA H100, NVIDIA H200, and AMD MI300X GPUs. Our results reveal that no universal optimal power cap exists, as the efficiency peak varies across application types and GPU architectures. Interestingly, the two NVIDIA GPUs which mainly differ in their HBM configuration show qualitatively different performance-energy trade-offs. The developed benchmarking framework will be released as a public tool.
title AI Application Benchmarking: Power-Aware Performance Analysis for Vision and Language Models
topic Performance
url https://arxiv.org/abs/2603.16164