AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zikun, Chen, Zhuofu, Delacourt, Remi, Oliaro, Gabriele, Wang, Zeyu, Chen, Qinghan, Lin, Shuhuai, Yang, April, Zhang, Zhihao, Chen, Zhuoming, Lai, Sean, Cheng, Xinhao, Miao, Xupeng, Jia, Zhihao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
by: Oliaro, Gabriele, et al.
Published: (2024)
by: Oliaro, Gabriele, et al.
Published: (2024)
AdaSpec: Adaptive Speculative Decoding for Fast, SLO-Aware Large Language Model Serving
by: Huang, Kaiyu, et al.
Published: (2025)
by: Huang, Kaiyu, et al.
Published: (2025)
Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
by: Miao, Xupeng, et al.
Published: (2023)
by: Miao, Xupeng, et al.
Published: (2023)
SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification
by: Miao, Xupeng, et al.
Published: (2023)
by: Miao, Xupeng, et al.
Published: (2023)
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
by: Chen, Siyuan, et al.
Published: (2025)
by: Chen, Siyuan, et al.
Published: (2025)
PolyServe: Efficient Multi-SLO Serving at Scale
by: Zhu, Kan, et al.
Published: (2025)
by: Zhu, Kan, et al.
Published: (2025)
SuffixDecoding: Extreme Speculative Decoding for Emerging AI Applications
by: Oliaro, Gabriele, et al.
Published: (2024)
by: Oliaro, Gabriele, et al.
Published: (2024)
TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention
by: Yang, Lijie, et al.
Published: (2024)
by: Yang, Lijie, et al.
Published: (2024)
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
by: Shen, Haiying, et al.
Published: (2024)
by: Shen, Haiying, et al.
Published: (2024)
JITServe: SLO-aware LLM Serving with Imprecise Request Information
by: Zhang, Wei, et al.
Published: (2025)
by: Zhang, Wei, et al.
Published: (2025)
SLO-Guard: Crash-Aware, Budget-Consistent Autotuning for SLO-Constrained LLM Serving
by: Lysenstøen, Christian
Published: (2026)
by: Lysenstøen, Christian
Published: (2026)
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
by: Nie, Chengyi, et al.
Published: (2024)
by: Nie, Chengyi, et al.
Published: (2024)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
by: Mo, Zizhao, et al.
Published: (2026)
by: Mo, Zizhao, et al.
Published: (2026)
HFX: Joint Design of Algorithms and Systems for Multi-SLO Serving and Fast Scaling
by: Yousefijamarani, Zahra, et al.
Published: (2025)
by: Yousefijamarani, Zahra, et al.
Published: (2025)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
by: Hu, Jianmin, et al.
Published: (2025)
by: Hu, Jianmin, et al.
Published: (2025)
Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs
by: Mei, Yixuan, et al.
Published: (2026)
by: Mei, Yixuan, et al.
Published: (2026)
PATCHEDSERVE: A Patch Management Framework for SLO-Optimized Hybrid Resolution Diffusion Serving
by: Sun, Desen, et al.
Published: (2025)
by: Sun, Desen, et al.
Published: (2025)
Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding
by: Chen, Zhuoming, et al.
Published: (2024)
by: Chen, Zhuoming, et al.
Published: (2024)
SpecExec: Massively Parallel Speculative Decoding for Interactive LLM Inference on Consumer Devices
by: Svirschevski, Ruslan, et al.
Published: (2024)
by: Svirschevski, Ruslan, et al.
Published: (2024)
AccelGen: Heterogeneous SLO-Guaranteed High-Throughput LLM Inference Serving for Diverse Applications
by: Shen, Haiying, et al.
Published: (2025)
by: Shen, Haiying, et al.
Published: (2025)
SLO-aware GPU Frequency Scaling for Energy Efficient LLM Inference Serving
by: Kakolyris, Andreas Kosmas, et al.
Published: (2024)
by: Kakolyris, Andreas Kosmas, et al.
Published: (2024)
MACE: A Hybrid LLM Serving System with Colocated SLO-aware Continuous Retraining Alignment
by: Li, Yufei, et al.
Published: (2025)
by: Li, Yufei, et al.
Published: (2025)
Accelerating Retrieval-Augmented Language Model Serving with Speculation
by: Zhang, Zhihao, et al.
Published: (2024)
by: Zhang, Zhihao, et al.
Published: (2024)
CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters
by: Huang, Shaoyuan, et al.
Published: (2026)
by: Huang, Shaoyuan, et al.
Published: (2026)
WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching
by: Li, Xiangchen, et al.
Published: (2026)
by: Li, Xiangchen, et al.
Published: (2026)
Compass: SLO-aware Query Planner for Compound AI Serving at Scale
by: Liu, Banruo, et al.
Published: (2025)
by: Liu, Banruo, et al.
Published: (2025)
WWW.Serve: Interconnecting Global LLM Services through Decentralization
by: Wang, Huanyu, et al.
Published: (2026)
by: Wang, Huanyu, et al.
Published: (2026)
Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-Flow
by: Mei, Yixuan, et al.
Published: (2024)
by: Mei, Yixuan, et al.
Published: (2024)
Hummingbird: SLO-Oriented GPU Preemption at Microsecond-scale
by: Hu, Tiancheng, et al.
Published: (2026)
by: Hu, Tiancheng, et al.
Published: (2026)
Aragog: Just-in-Time Model Routing for Scalable Serving of Agentic Workflows
by: Dai, Yinwei, et al.
Published: (2025)
by: Dai, Yinwei, et al.
Published: (2025)
FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving
by: Chen, Wenyan, et al.
Published: (2026)
by: Chen, Wenyan, et al.
Published: (2026)
SLO-Aware Scheduling for Large Language Model Inferences
by: Huang, Jinqi, et al.
Published: (2025)
by: Huang, Jinqi, et al.
Published: (2025)
Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving
by: Shi, Xiaoxiang, et al.
Published: (2025)
by: Shi, Xiaoxiang, et al.
Published: (2025)
Nightjar: Dynamic Adaptive Speculative Decoding for Large Language Models Serving
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
CAT Pruning: Cluster-Aware Token Pruning For Text-to-Image Diffusion Models
by: Cheng, Xinle, et al.
Published: (2025)
by: Cheng, Xinle, et al.
Published: (2025)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
by: Chen, Jiabin, et al.
Published: (2024)
by: Chen, Jiabin, et al.
Published: (2024)
SLO-Aware Compute Resource Allocation for Prefill-Decode Disaggregated LLM Inference
by: Li, Luchang, et al.
Published: (2026)
by: Li, Luchang, et al.
Published: (2026)
TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding
by: Sun, Hanshi, et al.
Published: (2024)
by: Sun, Hanshi, et al.
Published: (2024)
Conveyor: Efficient Tool-aware LLM Serving with Tool Partial Execution
by: Xu, Yechen, et al.
Published: (2024)
by: Xu, Yechen, et al.
Published: (2024)
MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding
by: Sadhukhan, Ranajoy, et al.
Published: (2024)
by: Sadhukhan, Ranajoy, et al.
Published: (2024)
Similar Items
-
FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
by: Oliaro, Gabriele, et al.
Published: (2024) -
AdaSpec: Adaptive Speculative Decoding for Fast, SLO-Aware Large Language Model Serving
by: Huang, Kaiyu, et al.
Published: (2025) -
Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
by: Miao, Xupeng, et al.
Published: (2023) -
SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification
by: Miao, Xupeng, et al.
Published: (2023) -
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
by: Chen, Siyuan, et al.
Published: (2025)