Prism: Unleashing GPU Sharing for Cost-Efficient Multi-LLM Serving

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yu, Shan, Xing, Jiarong, Qiao, Yifan, Ma, Mingyuan, Li, Yangmin, Wang, Yang, Yang, Shuo, Xie, Zhiqiang, Cao, Shiyi, Bao, Ke, Stoica, Ion, Xu, Harry, Sheng, Ying
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915285701754880
author Yu, Shan
Xing, Jiarong
Qiao, Yifan
Ma, Mingyuan
Li, Yangmin
Wang, Yang
Yang, Shuo
Xie, Zhiqiang
Cao, Shiyi
Bao, Ke
Stoica, Ion
Xu, Harry
Sheng, Ying
author_facet Yu, Shan
Xing, Jiarong
Qiao, Yifan
Ma, Mingyuan
Li, Yangmin
Wang, Yang
Yang, Shuo
Xie, Zhiqiang
Cao, Shiyi
Bao, Ke
Stoica, Ion
Xu, Harry
Sheng, Ying
contents Serving large language models (LLMs) is expensive, especially for providers hosting many models, making cost reduction essential. The unique workload patterns of serving multiple LLMs (i.e., multi-LLM serving) create new opportunities and challenges for this task. The long-tail popularity of models and their long idle periods present opportunities to improve utilization through GPU sharing. However, existing GPU sharing systems lack the ability to adjust their resource allocation and sharing policies at runtime, making them ineffective at meeting latency service-level objectives (SLOs) under rapidly fluctuating workloads. This paper presents Prism, a multi-LLM serving system that unleashes the full potential of GPU sharing to achieve both cost efficiency and SLO attainment. At its core, Prism tackles a key limitation of existing systems$\unicode{x2014}$the lack of $\textit{cross-model memory coordination}$, which is essential for flexibly sharing GPU memory across models under dynamic workloads. Prism achieves this with two key designs. First, it supports on-demand memory allocation by dynamically mapping physical to virtual memory pages, allowing flexible memory redistribution among models that space- and time-share a GPU. Second, it improves memory efficiency through a two-level scheduling policy that dynamically adjusts sharing strategies based on models' runtime demands. Evaluations on real-world traces show that Prism achieves more than $2\times$ cost savings and $3.3\times$ SLO attainment compared to state-of-the-art systems.
format Preprint
id arxiv_https___arxiv_org_abs_2505_04021
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Prism: Unleashing GPU Sharing for Cost-Efficient Multi-LLM Serving
Yu, Shan
Xing, Jiarong
Qiao, Yifan
Ma, Mingyuan
Li, Yangmin
Wang, Yang
Yang, Shuo
Xie, Zhiqiang
Cao, Shiyi
Bao, Ke
Stoica, Ion
Xu, Harry
Sheng, Ying
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
Performance
Serving large language models (LLMs) is expensive, especially for providers hosting many models, making cost reduction essential. The unique workload patterns of serving multiple LLMs (i.e., multi-LLM serving) create new opportunities and challenges for this task. The long-tail popularity of models and their long idle periods present opportunities to improve utilization through GPU sharing. However, existing GPU sharing systems lack the ability to adjust their resource allocation and sharing policies at runtime, making them ineffective at meeting latency service-level objectives (SLOs) under rapidly fluctuating workloads. This paper presents Prism, a multi-LLM serving system that unleashes the full potential of GPU sharing to achieve both cost efficiency and SLO attainment. At its core, Prism tackles a key limitation of existing systems$\unicode{x2014}$the lack of $\textit{cross-model memory coordination}$, which is essential for flexibly sharing GPU memory across models under dynamic workloads. Prism achieves this with two key designs. First, it supports on-demand memory allocation by dynamically mapping physical to virtual memory pages, allowing flexible memory redistribution among models that space- and time-share a GPU. Second, it improves memory efficiency through a two-level scheduling policy that dynamically adjusts sharing strategies based on models' runtime demands. Evaluations on real-world traces show that Prism achieves more than $2\times$ cost savings and $3.3\times$ SLO attainment compared to state-of-the-art systems.
title Prism: Unleashing GPU Sharing for Cost-Efficient Multi-LLM Serving
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
Performance
url https://arxiv.org/abs/2505.04021