Leveraging Hardware Performance Counters for Predicting Workload Interference in Vector Supercomputers

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shubham, Takahashi, Keichi, Takizawa, Hiroyuki
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916450365603840
author Shubham
Takahashi, Keichi
Takizawa, Hiroyuki
author_facet Shubham
Takahashi, Keichi
Takizawa, Hiroyuki
contents In the rapidly evolving domain of high-performance computing (HPC), heterogeneous architectures such as the SX-Aurora TSUBASA (SX-AT) system architecture, which integrate diverse processor types, present both opportunities and challenges for optimizing resource utilization. This paper investigates workload interference within an SX-AT system, with a specific focus on resource contention between Vector Hosts (VHs) and Vector Engines (VEs). Through comprehensive empirical analysis, the study identifies key factors contributing to performance degradation, such as cache and memory bandwidth contention, when jobs with varying computational demands share resources. To address these issues, we develop a predictive model that leverages hardware performance counters (HCs) and machine learning (ML) algorithms to classify and predict workload interference. Our results demonstrate that the model accurately forecasts performance degradation, offering valuable insights for future research on optimizing job scheduling and resource allocation. This approach highlights the importance of adaptive resource management strategies in maintaining system efficiency and provides a foundation for future enhancements in heterogeneous supercomputing environments.
format Preprint
id arxiv_https___arxiv_org_abs_2410_18126
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Leveraging Hardware Performance Counters for Predicting Workload Interference in Vector Supercomputers
Shubham
Takahashi, Keichi
Takizawa, Hiroyuki
Distributed, Parallel, and Cluster Computing
In the rapidly evolving domain of high-performance computing (HPC), heterogeneous architectures such as the SX-Aurora TSUBASA (SX-AT) system architecture, which integrate diverse processor types, present both opportunities and challenges for optimizing resource utilization. This paper investigates workload interference within an SX-AT system, with a specific focus on resource contention between Vector Hosts (VHs) and Vector Engines (VEs). Through comprehensive empirical analysis, the study identifies key factors contributing to performance degradation, such as cache and memory bandwidth contention, when jobs with varying computational demands share resources. To address these issues, we develop a predictive model that leverages hardware performance counters (HCs) and machine learning (ML) algorithms to classify and predict workload interference. Our results demonstrate that the model accurately forecasts performance degradation, offering valuable insights for future research on optimizing job scheduling and resource allocation. This approach highlights the importance of adaptive resource management strategies in maintaining system efficiency and provides a foundation for future enhancements in heterogeneous supercomputing environments.
title Leveraging Hardware Performance Counters for Predicting Workload Interference in Vector Supercomputers
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2410.18126