Optimizing Hardware Resource Partitioning and Job Allocations on Modern GPUs under Power Caps

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Arima, Eishi, Kang, Minjoon, Saba, Issa, Weidendorfer, Josef, Trinitis, Carsten, Schulz, Martin
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909192401453056
author Arima, Eishi
Kang, Minjoon
Saba, Issa
Weidendorfer, Josef
Trinitis, Carsten
Schulz, Martin
author_facet Arima, Eishi
Kang, Minjoon
Saba, Issa
Weidendorfer, Josef
Trinitis, Carsten
Schulz, Martin
contents CPU-GPU heterogeneous systems are now commonly used in HPC (High-Performance Computing). However, improving the utilization and energy-efficiency of such systems is still one of the most critical issues. As one single program typically cannot fully utilize all resources within a node/chip, co-scheduling (or co-locating) multiple programs with complementary resource requirements is a promising solution. Meanwhile, as power consumption has become the first-class design constraint for HPC systems, such co-scheduling techniques should be well-tailored for power-constrained environments. To this end, the industry recently started supporting hardware-level resource partitioning features on modern GPUs for realizing efficient co-scheduling, which can operate with existing power capping features. For example, NVidia's MIG (Multi-Instance GPU) partitions one single GPU into multiple instances at the granularity of a GPC (Graphics Processing Cluster). In this paper, we explicitly target the combination of hardware-level GPU partitioning features and power capping for power-constrained HPC systems. We provide a systematic methodology to optimize the combination of chip partitioning, job allocations, as well as power capping based on our scalability/interference modeling while taking a variety of aspects into account, such as compute/memory intensity and utilization in heterogeneous computational resources (e.g., Tensor Cores). The experimental result indicates that our approach is successful in selecting a near optimal combination across multiple different workloads.
format Preprint
id arxiv_https___arxiv_org_abs_2405_03838
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Optimizing Hardware Resource Partitioning and Job Allocations on Modern GPUs under Power Caps
Arima, Eishi
Kang, Minjoon
Saba, Issa
Weidendorfer, Josef
Trinitis, Carsten
Schulz, Martin
Distributed, Parallel, and Cluster Computing
CPU-GPU heterogeneous systems are now commonly used in HPC (High-Performance Computing). However, improving the utilization and energy-efficiency of such systems is still one of the most critical issues. As one single program typically cannot fully utilize all resources within a node/chip, co-scheduling (or co-locating) multiple programs with complementary resource requirements is a promising solution. Meanwhile, as power consumption has become the first-class design constraint for HPC systems, such co-scheduling techniques should be well-tailored for power-constrained environments. To this end, the industry recently started supporting hardware-level resource partitioning features on modern GPUs for realizing efficient co-scheduling, which can operate with existing power capping features. For example, NVidia's MIG (Multi-Instance GPU) partitions one single GPU into multiple instances at the granularity of a GPC (Graphics Processing Cluster). In this paper, we explicitly target the combination of hardware-level GPU partitioning features and power capping for power-constrained HPC systems. We provide a systematic methodology to optimize the combination of chip partitioning, job allocations, as well as power capping based on our scalability/interference modeling while taking a variety of aspects into account, such as compute/memory intensity and utilization in heterogeneous computational resources (e.g., Tensor Cores). The experimental result indicates that our approach is successful in selecting a near optimal combination across multiple different workloads.
title Optimizing Hardware Resource Partitioning and Job Allocations on Modern GPUs under Power Caps
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2405.03838