Mitigating Shared Storage Congestion Using Control Theory

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Collignon, Thomas, Halitim, Kouds, Bleuse, Raphaël, Cerf, Sophie, Robu, Bogdan, Rutten, Éric, Seinturier, Lionel, van Kempen, Alexandre
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914165439856640
author Collignon, Thomas
Halitim, Kouds
Bleuse, Raphaël
Cerf, Sophie
Robu, Bogdan
Rutten, Éric
Seinturier, Lionel
van Kempen, Alexandre
author_facet Collignon, Thomas
Halitim, Kouds
Bleuse, Raphaël
Cerf, Sophie
Robu, Bogdan
Rutten, Éric
Seinturier, Lionel
van Kempen, Alexandre
contents Efficient data access in High-Performance Computing (HPC) systems is essential to the performance of intensive computing tasks. Traditional optimizations of the I/O stack aim to improve peak performance but are often workload specific and require deep expertise, making them difficult to generalize or re-use. In shared HPC environments, resource congestion can lead to unpredictable performance, causing slowdowns and timeouts. To address these challenges, we propose a self-adaptive approach based on Control Theory to dynamically regulate client-side I/O rates. Our approach leverages a small set of runtime system load metrics to reduce congestion and enhance performance stability. We implement a controller in a multi-node cluster and evaluate it on a real testbed under a representative workload. Experimental results demonstrate that our method effectively mitigates I/O congestion, reducing total runtime by up to 20% and lowering tail latency, while maintaining stable performance.
format Preprint
id arxiv_https___arxiv_org_abs_2511_16177
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mitigating Shared Storage Congestion Using Control Theory
Collignon, Thomas
Halitim, Kouds
Bleuse, Raphaël
Cerf, Sophie
Robu, Bogdan
Rutten, Éric
Seinturier, Lionel
van Kempen, Alexandre
Distributed, Parallel, and Cluster Computing
Hardware Architecture
Efficient data access in High-Performance Computing (HPC) systems is essential to the performance of intensive computing tasks. Traditional optimizations of the I/O stack aim to improve peak performance but are often workload specific and require deep expertise, making them difficult to generalize or re-use. In shared HPC environments, resource congestion can lead to unpredictable performance, causing slowdowns and timeouts. To address these challenges, we propose a self-adaptive approach based on Control Theory to dynamically regulate client-side I/O rates. Our approach leverages a small set of runtime system load metrics to reduce congestion and enhance performance stability. We implement a controller in a multi-node cluster and evaluate it on a real testbed under a representative workload. Experimental results demonstrate that our method effectively mitigates I/O congestion, reducing total runtime by up to 20% and lowering tail latency, while maintaining stable performance.
title Mitigating Shared Storage Congestion Using Control Theory
topic Distributed, Parallel, and Cluster Computing
Hardware Architecture
url https://arxiv.org/abs/2511.16177