Revisiting Service Level Objectives and System Level Metrics in Large Language Model Serving

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Zhibin, Li, Shipeng, Zhou, Yuhang, Li, Xue, Zhang, Zhonghui, Cam-Tu, Nguyen, Gu, Rong, Tian, Chen, Chen, Guihai, Zhong, Sheng
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912674408824832
author Wang, Zhibin
Li, Shipeng
Zhou, Yuhang
Li, Xue
Zhang, Zhonghui
Cam-Tu, Nguyen
Gu, Rong
Tian, Chen
Chen, Guihai
Zhong, Sheng
author_facet Wang, Zhibin
Li, Shipeng
Zhou, Yuhang
Li, Xue
Zhang, Zhonghui
Cam-Tu, Nguyen
Gu, Rong
Tian, Chen
Chen, Guihai
Zhong, Sheng
contents User experience is a critical factor Large Language Model (LLM) serving systems must consider, where service level objectives (SLOs) considering the experience of individual requests and system level metrics (SLMs) considering the overall system performance are two key performance measures. However, we observe two notable issues in existing metrics: 1) manually delaying the delivery of some tokens can improve SLOs, and 2) actively abandoning requests that do not meet SLOs can improve SLMs, both of which are counterintuitive. In this paper, we revisit SLOs and SLMs in LLM serving, and propose a new SLO that aligns with user experience. Based on the SLO, we propose a comprehensive metric framework called smooth goodput, which integrates SLOs and SLMs to reflect the nature of user experience in LLM serving. Through this unified framework, we reassess the performance of different LLM serving systems under multiple workloads. Evaluation results show that our metric framework provides a more comprehensive view of token delivery and request processing, and effectively captures the optimal point of user experience and system performance with different serving strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2410_14257
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Revisiting Service Level Objectives and System Level Metrics in Large Language Model Serving
Wang, Zhibin
Li, Shipeng
Zhou, Yuhang
Li, Xue
Zhang, Zhonghui
Cam-Tu, Nguyen
Gu, Rong
Tian, Chen
Chen, Guihai
Zhong, Sheng
Machine Learning
Artificial Intelligence
User experience is a critical factor Large Language Model (LLM) serving systems must consider, where service level objectives (SLOs) considering the experience of individual requests and system level metrics (SLMs) considering the overall system performance are two key performance measures. However, we observe two notable issues in existing metrics: 1) manually delaying the delivery of some tokens can improve SLOs, and 2) actively abandoning requests that do not meet SLOs can improve SLMs, both of which are counterintuitive. In this paper, we revisit SLOs and SLMs in LLM serving, and propose a new SLO that aligns with user experience. Based on the SLO, we propose a comprehensive metric framework called smooth goodput, which integrates SLOs and SLMs to reflect the nature of user experience in LLM serving. Through this unified framework, we reassess the performance of different LLM serving systems under multiple workloads. Evaluation results show that our metric framework provides a more comprehensive view of token delivery and request processing, and effectively captures the optimal point of user experience and system performance with different serving strategies.
title Revisiting Service Level Objectives and System Level Metrics in Large Language Model Serving
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2410.14257