AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866908368100130816 |
|---|---|
| author | Li, Zikun Chen, Zhuofu Delacourt, Remi Oliaro, Gabriele Wang, Zeyu Chen, Qinghan Lin, Shuhuai Yang, April Zhang, Zhihao Chen, Zhuoming Lai, Sean Cheng, Xinhao Miao, Xupeng Jia, Zhihao |
| author_facet | Li, Zikun Chen, Zhuofu Delacourt, Remi Oliaro, Gabriele Wang, Zeyu Chen, Qinghan Lin, Shuhuai Yang, April Zhang, Zhihao Chen, Zhuoming Lai, Sean Cheng, Xinhao Miao, Xupeng Jia, Zhihao |
| contents | Modern large language model (LLM) applications exhibit diverse service-level objectives (SLOs), from low-latency requirements in interactive coding assistants to more relaxed constraints in data wrangling tasks. Existing LLM serving systems, which rely on uniform batching and scheduling strategies, often fail to meet these heterogeneous SLOs concurrently. We present AdaServe, the first LLM serving system designed to support efficient multi-SLO serving through SLO-customized speculative decoding. AdaServe formulates multi-SLO serving as a constrained optimization problem and introduces a hardware-aware algorithm that constructs a speculation tree tailored to each request's latency target. It features a speculate-select-verify pipeline that enables fine-grained control over decoding speed while maximizing system throughput. AdaServe further adapts to workload variation by dynamically adjusting speculation parameters. Evaluations across diverse workloads show that AdaServe reduces SLO violations by up to 4.3$\times$ and improves goodput by up to 1.9$\times$ compared to the best performing baselines, highlighting its effectiveness in multi-SLO serving. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_12162 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding Li, Zikun Chen, Zhuofu Delacourt, Remi Oliaro, Gabriele Wang, Zeyu Chen, Qinghan Lin, Shuhuai Yang, April Zhang, Zhihao Chen, Zhuoming Lai, Sean Cheng, Xinhao Miao, Xupeng Jia, Zhihao Computation and Language Artificial Intelligence Distributed, Parallel, and Cluster Computing Machine Learning Modern large language model (LLM) applications exhibit diverse service-level objectives (SLOs), from low-latency requirements in interactive coding assistants to more relaxed constraints in data wrangling tasks. Existing LLM serving systems, which rely on uniform batching and scheduling strategies, often fail to meet these heterogeneous SLOs concurrently. We present AdaServe, the first LLM serving system designed to support efficient multi-SLO serving through SLO-customized speculative decoding. AdaServe formulates multi-SLO serving as a constrained optimization problem and introduces a hardware-aware algorithm that constructs a speculation tree tailored to each request's latency target. It features a speculate-select-verify pipeline that enables fine-grained control over decoding speed while maximizing system throughput. AdaServe further adapts to workload variation by dynamically adjusting speculation parameters. Evaluations across diverse workloads show that AdaServe reduces SLO violations by up to 4.3$\times$ and improves goodput by up to 1.9$\times$ compared to the best performing baselines, highlighting its effectiveness in multi-SLO serving. |
| title | AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding |
| topic | Computation and Language Artificial Intelligence Distributed, Parallel, and Cluster Computing Machine Learning |
| url | https://arxiv.org/abs/2501.12162 |