Efficient Serving of LLM Applications with Probabilistic Demand Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yifei, Gan, Zuo, Gan, Zhenghao, Wang, Weiye, Chen, Chen, Shan, Yizhou, Chen, Xusheng, Han, Zhenhua, Zhu, Yifei, Sun, Shixuan, Guo, Minyi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912436878049280
author Liu, Yifei
Gan, Zuo
Gan, Zhenghao
Wang, Weiye
Chen, Chen
Shan, Yizhou
Chen, Xusheng
Han, Zhenhua
Zhu, Yifei
Sun, Shixuan
Guo, Minyi
author_facet Liu, Yifei
Gan, Zuo
Gan, Zhenghao
Wang, Weiye
Chen, Chen
Shan, Yizhou
Chen, Xusheng
Han, Zhenhua
Zhu, Yifei
Sun, Shixuan
Guo, Minyi
contents Applications based on Large Language Models (LLMs) contains a series of tasks to address real-world problems with boosted capability, which have dynamic demand volumes on diverse backends. Existing serving systems treat the resource demands of LLM applications as a blackbox, compromising end-to-end efficiency due to improper queuing order and backend warm up latency. We find that the resource demands of LLM applications can be modeled in a general and accurate manner with Probabilistic Demand Graph (PDGraph). We then propose Hermes, which leverages PDGraph for efficient serving of LLM applications. Confronting probabilistic demand description, Hermes applies the Gittins policy to determine the scheduling order that can minimize the average application completion time. It also uses the PDGraph model to help prewarm cold backends at proper moments. Experiments with diverse LLM applications confirm that Hermes can effectively improve the application serving efficiency, reducing the average completion time by over 70% and the P95 completion time by over 80%.
format Preprint
id arxiv_https___arxiv_org_abs_2506_14851
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient Serving of LLM Applications with Probabilistic Demand Modeling
Liu, Yifei
Gan, Zuo
Gan, Zhenghao
Wang, Weiye
Chen, Chen
Shan, Yizhou
Chen, Xusheng
Han, Zhenhua
Zhu, Yifei
Sun, Shixuan
Guo, Minyi
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
Applications based on Large Language Models (LLMs) contains a series of tasks to address real-world problems with boosted capability, which have dynamic demand volumes on diverse backends. Existing serving systems treat the resource demands of LLM applications as a blackbox, compromising end-to-end efficiency due to improper queuing order and backend warm up latency. We find that the resource demands of LLM applications can be modeled in a general and accurate manner with Probabilistic Demand Graph (PDGraph). We then propose Hermes, which leverages PDGraph for efficient serving of LLM applications. Confronting probabilistic demand description, Hermes applies the Gittins policy to determine the scheduling order that can minimize the average application completion time. It also uses the PDGraph model to help prewarm cold backends at proper moments. Experiments with diverse LLM applications confirm that Hermes can effectively improve the application serving efficiency, reducing the average completion time by over 70% and the P95 completion time by over 80%.
title Efficient Serving of LLM Applications with Probabilistic Demand Modeling
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.14851