A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via Faro

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jeon, Beomyeol, Wang, Chen, Arroyo, Diana, Youssef, Alaa, Gupta, Indranil
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914959976300544
author Jeon, Beomyeol
Wang, Chen
Arroyo, Diana
Youssef, Alaa
Gupta, Indranil
author_facet Jeon, Beomyeol
Wang, Chen
Arroyo, Diana
Youssef, Alaa
Gupta, Indranil
contents This paper tackles the challenge of running multiple ML inference jobs (models) under time-varying workloads, on a constrained on-premises production cluster. Our system Faro takes in latency Service Level Objectives (SLOs) for each job, auto-distills them into utility functions, "sloppifies" these utility functions to make them amenable to mathematical optimization, automatically predicts workload via probabilistic prediction, and dynamically makes implicit cross-job resource allocations, in order to satisfy cluster-wide objectives, e.g., total utility, fairness, and other hybrid variants. A major challenge Faro tackles is that using precise utilities and high-fidelity predictors, can be too slow (and in a sense too precise!) for the fast adaptation we require. Faro's solution is to "sloppify" (relax) its multiple design components to achieve fast adaptation without overly degrading solution quality. Faro is implemented in a stack consisting of Ray Serve running atop a Kubernetes cluster. Trace-driven cluster deployments show that Faro achieves 2.3$\times$-23$\times$ lower SLO violations compared to state-of-the-art systems.
format Preprint
id arxiv_https___arxiv_org_abs_2409_19488
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via Faro
Jeon, Beomyeol
Wang, Chen
Arroyo, Diana
Youssef, Alaa
Gupta, Indranil
Distributed, Parallel, and Cluster Computing
This paper tackles the challenge of running multiple ML inference jobs (models) under time-varying workloads, on a constrained on-premises production cluster. Our system Faro takes in latency Service Level Objectives (SLOs) for each job, auto-distills them into utility functions, "sloppifies" these utility functions to make them amenable to mathematical optimization, automatically predicts workload via probabilistic prediction, and dynamically makes implicit cross-job resource allocations, in order to satisfy cluster-wide objectives, e.g., total utility, fairness, and other hybrid variants. A major challenge Faro tackles is that using precise utilities and high-fidelity predictors, can be too slow (and in a sense too precise!) for the fast adaptation we require. Faro's solution is to "sloppify" (relax) its multiple design components to achieve fast adaptation without overly degrading solution quality. Faro is implemented in a stack consisting of Ray Serve running atop a Kubernetes cluster. Trace-driven cluster deployments show that Faro achieves 2.3$\times$-23$\times$ lower SLO violations compared to state-of-the-art systems.
title A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via Faro
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2409.19488