vPALs: Towards Verified Performance-aware Learning System For Resource Management

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: He, Guoliang, Yeung, Gingfung, Ceesay, Sheriffo, Barker, Adam
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909160428273664
author He, Guoliang
Yeung, Gingfung
Ceesay, Sheriffo
Barker, Adam
author_facet He, Guoliang
Yeung, Gingfung
Ceesay, Sheriffo
Barker, Adam
contents Accurately predicting task performance at runtime in a cluster is advantageous for a resource management system to determine whether a task should be migrated due to performance degradation caused by interference. This is beneficial for both cluster operators and service owners. However, deploying performance prediction systems with learning methods requires sophisticated safeguard mechanisms due to the inherent stochastic and black-box natures of these models, such as Deep Neural Networks (DNNs). Vanilla Neural Networks (NNs) can be vulnerable to out-of-distribution data samples that can lead to sub-optimal decisions. To take a step towards a safe learning system in performance prediction, We propose vPALs that leverage well-correlated system metrics, and verification to produce safe performance prediction at runtime, providing an extra layer of safety to integrate learning techniques to cluster resource management systems. Our experiments show that vPALs can outperform vanilla NNs across our benchmark workload.
format Preprint
id arxiv_https___arxiv_org_abs_2404_03079
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle vPALs: Towards Verified Performance-aware Learning System For Resource Management
He, Guoliang
Yeung, Gingfung
Ceesay, Sheriffo
Barker, Adam
Distributed, Parallel, and Cluster Computing
Accurately predicting task performance at runtime in a cluster is advantageous for a resource management system to determine whether a task should be migrated due to performance degradation caused by interference. This is beneficial for both cluster operators and service owners. However, deploying performance prediction systems with learning methods requires sophisticated safeguard mechanisms due to the inherent stochastic and black-box natures of these models, such as Deep Neural Networks (DNNs). Vanilla Neural Networks (NNs) can be vulnerable to out-of-distribution data samples that can lead to sub-optimal decisions. To take a step towards a safe learning system in performance prediction, We propose vPALs that leverage well-correlated system metrics, and verification to produce safe performance prediction at runtime, providing an extra layer of safety to integrate learning techniques to cluster resource management systems. Our experiments show that vPALs can outperform vanilla NNs across our benchmark workload.
title vPALs: Towards Verified Performance-aware Learning System For Resource Management
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2404.03079