MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Duan, Jiangfei, Lu, Runyu, Duanmu, Haojie, Li, Xiuhong, Zhang, Xingcheng, Lin, Dahua, Stoica, Ion, Zhang, Hao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916284646555648
author Duan, Jiangfei
Lu, Runyu
Duanmu, Haojie
Li, Xiuhong
Zhang, Xingcheng
Lin, Dahua
Stoica, Ion
Zhang, Hao
author_facet Duan, Jiangfei
Lu, Runyu
Duanmu, Haojie
Li, Xiuhong
Zhang, Xingcheng
Lin, Dahua
Stoica, Ion
Zhang, Hao
contents Large language models (LLMs) have demonstrated remarkable performance, and organizations are racing to serve LLMs of varying sizes as endpoints for use-cases like chat, programming and search. However, efficiently serving multiple LLMs poses significant challenges for existing approaches due to varying popularity of LLMs. In the paper, we present MuxServe, a flexible spatial-temporal multiplexing system for efficient multiple LLM serving. The key insight behind is to colocate LLMs considering their popularity to multiplex memory resources, and leverage the characteristics of prefill and decoding phases to separate and flexibly colocate them to multiplex computation resources. MuxServe formally formulates the multiplexing problem, and proposes a novel placement algorithm and adaptive batch scheduling strategy to identify optimal colocations and maximize utilization. MuxServe designs a unified resource manager to enable flexible and efficient multiplexing. Evaluation results show that MuxServe can achieves up to $1.8\times$ higher throughput or processes $2.9\times$ more requests within $99\%$ SLO attainment. The code is available at: \url{https://github.com/hao-ai-lab/MuxServe}.
format Preprint
id arxiv_https___arxiv_org_abs_2404_02015
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
Duan, Jiangfei
Lu, Runyu
Duanmu, Haojie
Li, Xiuhong
Zhang, Xingcheng
Lin, Dahua
Stoica, Ion
Zhang, Hao
Distributed, Parallel, and Cluster Computing
Large language models (LLMs) have demonstrated remarkable performance, and organizations are racing to serve LLMs of varying sizes as endpoints for use-cases like chat, programming and search. However, efficiently serving multiple LLMs poses significant challenges for existing approaches due to varying popularity of LLMs. In the paper, we present MuxServe, a flexible spatial-temporal multiplexing system for efficient multiple LLM serving. The key insight behind is to colocate LLMs considering their popularity to multiplex memory resources, and leverage the characteristics of prefill and decoding phases to separate and flexibly colocate them to multiplex computation resources. MuxServe formally formulates the multiplexing problem, and proposes a novel placement algorithm and adaptive batch scheduling strategy to identify optimal colocations and maximize utilization. MuxServe designs a unified resource manager to enable flexible and efficient multiplexing. Evaluation results show that MuxServe can achieves up to $1.8\times$ higher throughput or processes $2.9\times$ more requests within $99\%$ SLO attainment. The code is available at: \url{https://github.com/hao-ai-lab/MuxServe}.
title MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2404.02015