StraightLine: An End-to-End Resource-Aware Scheduler for Machine Learning Application Requests

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ching, Cheng-Wei, Guan, Boyuan, Xu, Hailu, Hu, Liting
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912017118396416
author Ching, Cheng-Wei
Guan, Boyuan
Xu, Hailu
Hu, Liting
author_facet Ching, Cheng-Wei
Guan, Boyuan
Xu, Hailu
Hu, Liting
contents The life cycle of machine learning (ML) applications consists of two stages: model development and model deployment. However, traditional ML systems (e.g., training-specific or inference-specific systems) focus on one particular stage or phase of the life cycle of ML applications. These systems often aim at optimizing model training or accelerating model inference, and they frequently assume homogeneous infrastructure, which may not always reflect real-world scenarios that include cloud data centers, local servers, containers, and serverless platforms. We present StraightLine, an end-to-end resource-aware scheduler that schedules the optimal resources (e.g., container, virtual machine, or serverless) for different ML application requests in a hybrid infrastructure. The key innovation is an empirical dynamic placing algorithm that intelligently places requests based on their unique characteristics (e.g., request frequency, input data size, and data distribution). In contrast to existing ML systems, StraightLine offers end-to-end resource-aware placement, thereby it can significantly reduce response time and failure rate for model deployment when facing different computing resources in the hybrid infrastructure.
format Preprint
id arxiv_https___arxiv_org_abs_2407_18148
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle StraightLine: An End-to-End Resource-Aware Scheduler for Machine Learning Application Requests
Ching, Cheng-Wei
Guan, Boyuan
Xu, Hailu
Hu, Liting
Distributed, Parallel, and Cluster Computing
Machine Learning
The life cycle of machine learning (ML) applications consists of two stages: model development and model deployment. However, traditional ML systems (e.g., training-specific or inference-specific systems) focus on one particular stage or phase of the life cycle of ML applications. These systems often aim at optimizing model training or accelerating model inference, and they frequently assume homogeneous infrastructure, which may not always reflect real-world scenarios that include cloud data centers, local servers, containers, and serverless platforms. We present StraightLine, an end-to-end resource-aware scheduler that schedules the optimal resources (e.g., container, virtual machine, or serverless) for different ML application requests in a hybrid infrastructure. The key innovation is an empirical dynamic placing algorithm that intelligently places requests based on their unique characteristics (e.g., request frequency, input data size, and data distribution). In contrast to existing ML systems, StraightLine offers end-to-end resource-aware placement, thereby it can significantly reduce response time and failure rate for model deployment when facing different computing resources in the hybrid infrastructure.
title StraightLine: An End-to-End Resource-Aware Scheduler for Machine Learning Application Requests
topic Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2407.18148