AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Zikun, Chen, Zhuofu, Delacourt, Remi, Oliaro, Gabriele, Wang, Zeyu, Chen, Qinghan, Lin, Shuhuai, Yang, April, Zhang, Zhihao, Chen, Zhuoming, Lai, Sean, Cheng, Xinhao, Miao, Xupeng, Jia, Zhihao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908368100130816
author Li, Zikun
Chen, Zhuofu
Delacourt, Remi
Oliaro, Gabriele
Wang, Zeyu
Chen, Qinghan
Lin, Shuhuai
Yang, April
Zhang, Zhihao
Chen, Zhuoming
Lai, Sean
Cheng, Xinhao
Miao, Xupeng
Jia, Zhihao
author_facet Li, Zikun
Chen, Zhuofu
Delacourt, Remi
Oliaro, Gabriele
Wang, Zeyu
Chen, Qinghan
Lin, Shuhuai
Yang, April
Zhang, Zhihao
Chen, Zhuoming
Lai, Sean
Cheng, Xinhao
Miao, Xupeng
Jia, Zhihao
contents Modern large language model (LLM) applications exhibit diverse service-level objectives (SLOs), from low-latency requirements in interactive coding assistants to more relaxed constraints in data wrangling tasks. Existing LLM serving systems, which rely on uniform batching and scheduling strategies, often fail to meet these heterogeneous SLOs concurrently. We present AdaServe, the first LLM serving system designed to support efficient multi-SLO serving through SLO-customized speculative decoding. AdaServe formulates multi-SLO serving as a constrained optimization problem and introduces a hardware-aware algorithm that constructs a speculation tree tailored to each request's latency target. It features a speculate-select-verify pipeline that enables fine-grained control over decoding speed while maximizing system throughput. AdaServe further adapts to workload variation by dynamically adjusting speculation parameters. Evaluations across diverse workloads show that AdaServe reduces SLO violations by up to 4.3$\times$ and improves goodput by up to 1.9$\times$ compared to the best performing baselines, highlighting its effectiveness in multi-SLO serving.
format Preprint
id arxiv_https___arxiv_org_abs_2501_12162
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
Li, Zikun
Chen, Zhuofu
Delacourt, Remi
Oliaro, Gabriele
Wang, Zeyu
Chen, Qinghan
Lin, Shuhuai
Yang, April
Zhang, Zhihao
Chen, Zhuoming
Lai, Sean
Cheng, Xinhao
Miao, Xupeng
Jia, Zhihao
Computation and Language
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Machine Learning
Modern large language model (LLM) applications exhibit diverse service-level objectives (SLOs), from low-latency requirements in interactive coding assistants to more relaxed constraints in data wrangling tasks. Existing LLM serving systems, which rely on uniform batching and scheduling strategies, often fail to meet these heterogeneous SLOs concurrently. We present AdaServe, the first LLM serving system designed to support efficient multi-SLO serving through SLO-customized speculative decoding. AdaServe formulates multi-SLO serving as a constrained optimization problem and introduces a hardware-aware algorithm that constructs a speculation tree tailored to each request's latency target. It features a speculate-select-verify pipeline that enables fine-grained control over decoding speed while maximizing system throughput. AdaServe further adapts to workload variation by dynamically adjusting speculation parameters. Evaluations across diverse workloads show that AdaServe reduces SLO violations by up to 4.3$\times$ and improves goodput by up to 1.9$\times$ compared to the best performing baselines, highlighting its effectiveness in multi-SLO serving.
title AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
topic Computation and Language
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2501.12162