IPA: Inference Pipeline Adaptation to Achieve High Accuracy and Cost-Efficiency

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ghafouri, Saeid, Razavi, Kamran, Salmani, Mehran, Sanaee, Alireza, Lorido-Botran, Tania, Wang, Lin, Doyle, Joseph, Jamshidi, Pooyan
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910459410513920
author Ghafouri, Saeid
Razavi, Kamran
Salmani, Mehran
Sanaee, Alireza
Lorido-Botran, Tania
Wang, Lin
Doyle, Joseph
Jamshidi, Pooyan
author_facet Ghafouri, Saeid
Razavi, Kamran
Salmani, Mehran
Sanaee, Alireza
Lorido-Botran, Tania
Wang, Lin
Doyle, Joseph
Jamshidi, Pooyan
contents Efficiently optimizing multi-model inference pipelines for fast, accurate, and cost-effective inference is a crucial challenge in machine learning production systems, given their tight end-to-end latency requirements. To simplify the exploration of the vast and intricate trade-off space of latency, accuracy, and cost in inference pipelines, providers frequently opt to consider one of them. However, the challenge lies in reconciling latency, accuracy, and cost trade-offs. To address this challenge and propose a solution to efficiently manage model variants in inference pipelines, we present IPA, an online deep learning Inference Pipeline Adaptation system that efficiently leverages model variants for each deep learning task. Model variants are different versions of pre-trained models for the same deep learning task with variations in resource requirements, latency, and accuracy. IPA dynamically configures batch size, replication, and model variants to optimize accuracy, minimize costs, and meet user-defined latency Service Level Agreements (SLAs) using Integer Programming. It supports multi-objective settings for achieving different trade-offs between accuracy and cost objectives while remaining adaptable to varying workloads and dynamic traffic patterns. Navigating a wider variety of configurations allows \namex{} to achieve better trade-offs between cost and accuracy objectives compared to existing methods. Extensive experiments in a Kubernetes implementation with five real-world inference pipelines demonstrate that IPA improves end-to-end accuracy by up to 21% with a minimal cost increase. The code and data for replications are available at https://github.com/reconfigurable-ml-pipeline/ipa.
format Preprint
id arxiv_https___arxiv_org_abs_2308_12871
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle IPA: Inference Pipeline Adaptation to Achieve High Accuracy and Cost-Efficiency
Ghafouri, Saeid
Razavi, Kamran
Salmani, Mehran
Sanaee, Alireza
Lorido-Botran, Tania
Wang, Lin
Doyle, Joseph
Jamshidi, Pooyan
Distributed, Parallel, and Cluster Computing
Machine Learning
Performance
Efficiently optimizing multi-model inference pipelines for fast, accurate, and cost-effective inference is a crucial challenge in machine learning production systems, given their tight end-to-end latency requirements. To simplify the exploration of the vast and intricate trade-off space of latency, accuracy, and cost in inference pipelines, providers frequently opt to consider one of them. However, the challenge lies in reconciling latency, accuracy, and cost trade-offs. To address this challenge and propose a solution to efficiently manage model variants in inference pipelines, we present IPA, an online deep learning Inference Pipeline Adaptation system that efficiently leverages model variants for each deep learning task. Model variants are different versions of pre-trained models for the same deep learning task with variations in resource requirements, latency, and accuracy. IPA dynamically configures batch size, replication, and model variants to optimize accuracy, minimize costs, and meet user-defined latency Service Level Agreements (SLAs) using Integer Programming. It supports multi-objective settings for achieving different trade-offs between accuracy and cost objectives while remaining adaptable to varying workloads and dynamic traffic patterns. Navigating a wider variety of configurations allows \namex{} to achieve better trade-offs between cost and accuracy objectives compared to existing methods. Extensive experiments in a Kubernetes implementation with five real-world inference pipelines demonstrate that IPA improves end-to-end accuracy by up to 21% with a minimal cost increase. The code and data for replications are available at https://github.com/reconfigurable-ml-pipeline/ipa.
title IPA: Inference Pipeline Adaptation to Achieve High Accuracy and Cost-Efficiency
topic Distributed, Parallel, and Cluster Computing
Machine Learning
Performance
url https://arxiv.org/abs/2308.12871