SLOs-Serve: Optimized Serving of Multi-SLO LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Siyuan, Jia, Zhipeng, Khan, Samira, Krishnamurthy, Arvind, Gibbons, Phillip B. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PolyServe: Efficient Multi-SLO Serving at Scale
von: Zhu, Kan, et al.
Veröffentlicht: (2025)
von: Zhu, Kan, et al.
Veröffentlicht: (2025)
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
von: Li, Zikun, et al.
Veröffentlicht: (2025)
von: Li, Zikun, et al.
Veröffentlicht: (2025)
Symphony: Optimized DNN Model Serving using Deferred Batch Scheduling
von: Chen, Lequn, et al.
Veröffentlicht: (2023)
von: Chen, Lequn, et al.
Veröffentlicht: (2023)
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
von: Shen, Haiying, et al.
Veröffentlicht: (2024)
von: Shen, Haiying, et al.
Veröffentlicht: (2024)
Sponge: Inference Serving with Dynamic SLOs Using In-Place Vertical Scaling
von: Razavi, Kamran, et al.
Veröffentlicht: (2024)
von: Razavi, Kamran, et al.
Veröffentlicht: (2024)
JITServe: SLO-aware LLM Serving with Imprecise Request Information
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
von: Oliaro, Gabriele, et al.
Veröffentlicht: (2024)
von: Oliaro, Gabriele, et al.
Veröffentlicht: (2024)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
von: Hu, Jianmin, et al.
Veröffentlicht: (2025)
von: Hu, Jianmin, et al.
Veröffentlicht: (2025)
WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
TetriServe: Efficient DiT Serving for Heterogeneous Image Generation
von: Lu, Runyu, et al.
Veröffentlicht: (2025)
von: Lu, Runyu, et al.
Veröffentlicht: (2025)
CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters
von: Huang, Shaoyuan, et al.
Veröffentlicht: (2026)
von: Huang, Shaoyuan, et al.
Veröffentlicht: (2026)
CascadeServe: Unlocking Model Cascades for Inference Serving
von: Kossmann, Ferdi, et al.
Veröffentlicht: (2024)
von: Kossmann, Ferdi, et al.
Veröffentlicht: (2024)
Cortex: Workflow-Aware Resource Pooling and Scheduling for Agentic Serving
von: Pagonas, Nikos, et al.
Veröffentlicht: (2025)
von: Pagonas, Nikos, et al.
Veröffentlicht: (2025)
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
von: Nie, Chengyi, et al.
Veröffentlicht: (2024)
von: Nie, Chengyi, et al.
Veröffentlicht: (2024)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
von: Ye, Zihao, et al.
Veröffentlicht: (2025)
von: Ye, Zihao, et al.
Veröffentlicht: (2025)
PATCHEDSERVE: A Patch Management Framework for SLO-Optimized Hybrid Resolution Diffusion Serving
von: Sun, Desen, et al.
Veröffentlicht: (2025)
von: Sun, Desen, et al.
Veröffentlicht: (2025)
ElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal Parallelism
von: Liu, Zedong, et al.
Veröffentlicht: (2025)
von: Liu, Zedong, et al.
Veröffentlicht: (2025)
DeltaZip: Efficient Serving of Multiple Full-Model-Tuned LLMs
von: Yao, Xiaozhe, et al.
Veröffentlicht: (2023)
von: Yao, Xiaozhe, et al.
Veröffentlicht: (2023)
ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving
von: Qiao, Yifan, et al.
Veröffentlicht: (2024)
von: Qiao, Yifan, et al.
Veröffentlicht: (2024)
LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism
von: Wu, Bingyang, et al.
Veröffentlicht: (2024)
von: Wu, Bingyang, et al.
Veröffentlicht: (2024)
MUSE: Multi-Tenant Model Serving With Seamless Model Updates
von: Correia, Cláudio, et al.
Veröffentlicht: (2026)
von: Correia, Cláudio, et al.
Veröffentlicht: (2026)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
von: Mo, Zizhao, et al.
Veröffentlicht: (2026)
von: Mo, Zizhao, et al.
Veröffentlicht: (2026)
FineServe: Precision-Aware KV Slab and Two-Level Scheduling for Heterogeneous Precision LLM Serving
von: Bin, Kyungmin, et al.
Veröffentlicht: (2025)
von: Bin, Kyungmin, et al.
Veröffentlicht: (2025)
EdgeServe: A Streaming System for Decentralized Model Serving
von: Shaowang, Ted, et al.
Veröffentlicht: (2023)
von: Shaowang, Ted, et al.
Veröffentlicht: (2023)
Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving
von: Shi, Xiaoxiang, et al.
Veröffentlicht: (2025)
von: Shi, Xiaoxiang, et al.
Veröffentlicht: (2025)
SLO-aware GPU Frequency Scaling for Energy Efficient LLM Inference Serving
von: Kakolyris, Andreas Kosmas, et al.
Veröffentlicht: (2024)
von: Kakolyris, Andreas Kosmas, et al.
Veröffentlicht: (2024)
QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration
von: Imani, HamidReza, et al.
Veröffentlicht: (2025)
von: Imani, HamidReza, et al.
Veröffentlicht: (2025)
OMEGA: A Low-Latency GNN Serving System for Large Graphs
von: Kim, Geon-Woo, et al.
Veröffentlicht: (2025)
von: Kim, Geon-Woo, et al.
Veröffentlicht: (2025)
VoxServe: Streaming-Centric Serving System for Speech Language Models
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2026)
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2026)
BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures
von: Hu, Xiannan, et al.
Veröffentlicht: (2025)
von: Hu, Xiannan, et al.
Veröffentlicht: (2025)
ViBE: Co-Optimizing Workload Skew and Hardware Variability for MoE Serving
von: Go, Seokjin, et al.
Veröffentlicht: (2026)
von: Go, Seokjin, et al.
Veröffentlicht: (2026)
MACE: A Hybrid LLM Serving System with Colocated SLO-aware Continuous Retraining Alignment
von: Li, Yufei, et al.
Veröffentlicht: (2025)
von: Li, Yufei, et al.
Veröffentlicht: (2025)
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
von: Su, Zhaoyuan, et al.
Veröffentlicht: (2025)
von: Su, Zhaoyuan, et al.
Veröffentlicht: (2025)
Locality-aware Fair Scheduling in LLM Serving
von: Cao, Shiyi, et al.
Veröffentlicht: (2025)
von: Cao, Shiyi, et al.
Veröffentlicht: (2025)
Towards Sustainable Large Language Model Serving
von: Nguyen, Sophia, et al.
Veröffentlicht: (2024)
von: Nguyen, Sophia, et al.
Veröffentlicht: (2024)
Stateful Large Language Model Serving with Pensieve
von: Yu, Lingfan, et al.
Veröffentlicht: (2023)
von: Yu, Lingfan, et al.
Veröffentlicht: (2023)
AIConfigurator: Lightning-Fast Configuration Optimization for Multi-Framework LLM Serving
von: Xu, Tianhao, et al.
Veröffentlicht: (2026)
von: Xu, Tianhao, et al.
Veröffentlicht: (2026)
MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing
von: Go, Seokjin, et al.
Veröffentlicht: (2025)
von: Go, Seokjin, et al.
Veröffentlicht: (2025)
P/D-Serve: Serving Disaggregated Large Language Model at Scale
von: Jin, Yibo, et al.
Veröffentlicht: (2024)
von: Jin, Yibo, et al.
Veröffentlicht: (2024)
HFX: Joint Design of Algorithms and Systems for Multi-SLO Serving and Fast Scaling
von: Yousefijamarani, Zahra, et al.
Veröffentlicht: (2025)
von: Yousefijamarani, Zahra, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
PolyServe: Efficient Multi-SLO Serving at Scale
von: Zhu, Kan, et al.
Veröffentlicht: (2025) -
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
von: Li, Zikun, et al.
Veröffentlicht: (2025) -
Symphony: Optimized DNN Model Serving using Deferred Batch Scheduling
von: Chen, Lequn, et al.
Veröffentlicht: (2023) -
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
von: Shen, Haiying, et al.
Veröffentlicht: (2024) -
Sponge: Inference Serving with Dynamic SLOs Using In-Place Vertical Scaling
von: Razavi, Kamran, et al.
Veröffentlicht: (2024)