ScaleLLM: A Resource-Frugal LLM Serving Framework by Optimizing End-to-End Efficiency

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yao, Yuhang, Jin, Han, Shah, Alay Dilipbhai, Han, Shanshan, Hu, Zijian, Ran, Yide, Stripelis, Dimitris, Xu, Zhaozhuo, Avestimehr, Salman, He, Chaoyang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912021868445696
author Yao, Yuhang
Jin, Han
Shah, Alay Dilipbhai
Han, Shanshan
Hu, Zijian
Ran, Yide
Stripelis, Dimitris
Xu, Zhaozhuo
Avestimehr, Salman
He, Chaoyang
author_facet Yao, Yuhang
Jin, Han
Shah, Alay Dilipbhai
Han, Shanshan
Hu, Zijian
Ran, Yide
Stripelis, Dimitris
Xu, Zhaozhuo
Avestimehr, Salman
He, Chaoyang
contents Large language models (LLMs) have surged in popularity and are extensively used in commercial applications, where the efficiency of model serving is crucial for the user experience. Most current research focuses on optimizing individual sub-procedures, e.g. local inference and communication, however, there is no comprehensive framework that provides a holistic system view for optimizing LLM serving in an end-to-end manner. In this work, we conduct a detailed analysis to identify major bottlenecks that impact end-to-end latency in LLM serving systems. Our analysis reveals that a comprehensive LLM serving endpoint must address a series of efficiency bottlenecks that extend beyond LLM inference. We then propose ScaleLLM, an optimized system for resource-efficient LLM serving. Our extensive experiments reveal that with 64 concurrent requests, ScaleLLM achieves a 4.3x speed up over vLLM and outperforms state-of-the-arts with 1.5x higher throughput.
format Preprint
id arxiv_https___arxiv_org_abs_2408_00008
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ScaleLLM: A Resource-Frugal LLM Serving Framework by Optimizing End-to-End Efficiency
Yao, Yuhang
Jin, Han
Shah, Alay Dilipbhai
Han, Shanshan
Hu, Zijian
Ran, Yide
Stripelis, Dimitris
Xu, Zhaozhuo
Avestimehr, Salman
He, Chaoyang
Distributed, Parallel, and Cluster Computing
Machine Learning
Large language models (LLMs) have surged in popularity and are extensively used in commercial applications, where the efficiency of model serving is crucial for the user experience. Most current research focuses on optimizing individual sub-procedures, e.g. local inference and communication, however, there is no comprehensive framework that provides a holistic system view for optimizing LLM serving in an end-to-end manner. In this work, we conduct a detailed analysis to identify major bottlenecks that impact end-to-end latency in LLM serving systems. Our analysis reveals that a comprehensive LLM serving endpoint must address a series of efficiency bottlenecks that extend beyond LLM inference. We then propose ScaleLLM, an optimized system for resource-efficient LLM serving. Our extensive experiments reveal that with 64 concurrent requests, ScaleLLM achieves a 4.3x speed up over vLLM and outperforms state-of-the-arts with 1.5x higher throughput.
title ScaleLLM: A Resource-Frugal LLM Serving Framework by Optimizing End-to-End Efficiency
topic Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2408.00008