ServerlessLLM: Low-Latency Serverless Inference for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fu, Yao, Xue, Leyang, Huang, Yeqi, Brabete, Andrei-Octavian, Ustiugov, Dmitrii, Patel, Yuvraj, Mai, Luo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910540625870848
author Fu, Yao
Xue, Leyang
Huang, Yeqi
Brabete, Andrei-Octavian
Ustiugov, Dmitrii
Patel, Yuvraj
Mai, Luo
author_facet Fu, Yao
Xue, Leyang
Huang, Yeqi
Brabete, Andrei-Octavian
Ustiugov, Dmitrii
Patel, Yuvraj
Mai, Luo
contents This paper presents ServerlessLLM, a distributed system designed to support low-latency serverless inference for Large Language Models (LLMs). By harnessing the substantial near-GPU storage and memory capacities of inference servers, ServerlessLLM achieves effective local checkpoint storage, minimizing the need for remote checkpoint downloads and ensuring efficient checkpoint loading. The design of ServerlessLLM features three core contributions: (i) \emph{fast multi-tier checkpoint loading}, featuring a new loading-optimized checkpoint format and a multi-tier loading system, fully utilizing the bandwidth of complex storage hierarchies on GPU servers; (ii) \emph{efficient live migration of LLM inference}, which enables newly initiated inferences to capitalize on local checkpoint storage while ensuring minimal user interruption; and (iii) \emph{startup-time-optimized model scheduling}, which assesses the locality statuses of checkpoints on each server and schedules the model onto servers that minimize the time to start the inference. Comprehensive evaluations, including microbenchmarks and real-world scenarios, demonstrate that ServerlessLLM dramatically outperforms state-of-the-art serverless systems, reducing latency by 10 - 200X across various LLM inference workloads.
format Preprint
id arxiv_https___arxiv_org_abs_2401_14351
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ServerlessLLM: Low-Latency Serverless Inference for Large Language Models
Fu, Yao
Xue, Leyang
Huang, Yeqi
Brabete, Andrei-Octavian
Ustiugov, Dmitrii
Patel, Yuvraj
Mai, Luo
Machine Learning
Distributed, Parallel, and Cluster Computing
This paper presents ServerlessLLM, a distributed system designed to support low-latency serverless inference for Large Language Models (LLMs). By harnessing the substantial near-GPU storage and memory capacities of inference servers, ServerlessLLM achieves effective local checkpoint storage, minimizing the need for remote checkpoint downloads and ensuring efficient checkpoint loading. The design of ServerlessLLM features three core contributions: (i) \emph{fast multi-tier checkpoint loading}, featuring a new loading-optimized checkpoint format and a multi-tier loading system, fully utilizing the bandwidth of complex storage hierarchies on GPU servers; (ii) \emph{efficient live migration of LLM inference}, which enables newly initiated inferences to capitalize on local checkpoint storage while ensuring minimal user interruption; and (iii) \emph{startup-time-optimized model scheduling}, which assesses the locality statuses of checkpoints on each server and schedules the model onto servers that minimize the time to start the inference. Comprehensive evaluations, including microbenchmarks and real-world scenarios, demonstrate that ServerlessLLM dramatically outperforms state-of-the-art serverless systems, reducing latency by 10 - 200X across various LLM inference workloads.
title ServerlessLLM: Low-Latency Serverless Inference for Large Language Models
topic Machine Learning
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2401.14351