Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
Fuente:
arXiv
Saved in:
| Main Authors: | Nie, Chengyi, Fonseca, Rodrigo, Liu, Zhenhua |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
by: Shen, Haiying, et al.
Published: (2024)
by: Shen, Haiying, et al.
Published: (2024)
Training DNN Models over Heterogeneous Clusters with Optimal Performance
by: Nie, Chengyi, et al.
Published: (2024)
by: Nie, Chengyi, et al.
Published: (2024)
DualScale: Energy-Efficient Disaggregated LLM Serving via Phase-Aware Placement and DVFS
by: Basit, Omar, et al.
Published: (2026)
by: Basit, Omar, et al.
Published: (2026)
PolyServe: Efficient Multi-SLO Serving at Scale
by: Zhu, Kan, et al.
Published: (2025)
by: Zhu, Kan, et al.
Published: (2025)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
by: Hu, Jianmin, et al.
Published: (2025)
by: Hu, Jianmin, et al.
Published: (2025)
HFX: Joint Design of Algorithms and Systems for Multi-SLO Serving and Fast Scaling
by: Yousefijamarani, Zahra, et al.
Published: (2025)
by: Yousefijamarani, Zahra, et al.
Published: (2025)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
by: Mo, Zizhao, et al.
Published: (2026)
by: Mo, Zizhao, et al.
Published: (2026)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
by: Wang, Qipeng
Published: (2026)
by: Wang, Qipeng
Published: (2026)
PromptTuner: SLO-Aware Elastic System for LLM Prompt Tuning
by: Gao, Wei, et al.
Published: (2026)
by: Gao, Wei, et al.
Published: (2026)
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
by: Chen, Siyuan, et al.
Published: (2025)
by: Chen, Siyuan, et al.
Published: (2025)
PATCHEDSERVE: A Patch Management Framework for SLO-Optimized Hybrid Resolution Diffusion Serving
by: Sun, Desen, et al.
Published: (2025)
by: Sun, Desen, et al.
Published: (2025)
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
by: Jaiswal, Shashwat, et al.
Published: (2025)
by: Jaiswal, Shashwat, et al.
Published: (2025)
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
by: Li, Zikun, et al.
Published: (2025)
by: Li, Zikun, et al.
Published: (2025)
CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters
by: Huang, Shaoyuan, et al.
Published: (2026)
by: Huang, Shaoyuan, et al.
Published: (2026)
SLO-Aware Scheduling for Large Language Model Inferences
by: Huang, Jinqi, et al.
Published: (2025)
by: Huang, Jinqi, et al.
Published: (2025)
EcoServe: Designing Carbon-Aware AI Inference Systems
by: Li, Yueying, et al.
Published: (2025)
by: Li, Yueying, et al.
Published: (2025)
SLO-Aware Task Offloading within Collaborative Vehicle Platoons
by: Sedlak, Boris, et al.
Published: (2024)
by: Sedlak, Boris, et al.
Published: (2024)
JITServe: SLO-aware LLM Serving with Imprecise Request Information
by: Zhang, Wei, et al.
Published: (2025)
by: Zhang, Wei, et al.
Published: (2025)
SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
by: Cheng, Ke, et al.
Published: (2024)
by: Cheng, Ke, et al.
Published: (2024)
HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
by: Lou, Chiheng, et al.
Published: (2025)
by: Lou, Chiheng, et al.
Published: (2025)
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
by: Chow, Will
Published: (2025)
by: Chow, Will
Published: (2025)
FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
by: Oliaro, Gabriele, et al.
Published: (2024)
by: Oliaro, Gabriele, et al.
Published: (2024)
DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling
by: Ahmad, Sohaib, et al.
Published: (2024)
by: Ahmad, Sohaib, et al.
Published: (2024)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
by: Li, Suyi, et al.
Published: (2024)
by: Li, Suyi, et al.
Published: (2024)
TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity
by: Lai, Ruiqi, et al.
Published: (2025)
by: Lai, Ruiqi, et al.
Published: (2025)
EdgeServing: Deadline-Aware Multi-DNN Serving at the Edge
by: Cao, Jiahe, et al.
Published: (2026)
by: Cao, Jiahe, et al.
Published: (2026)
DeepServe: Serverless Large Language Model Serving at Scale
by: Hu, Junhao, et al.
Published: (2025)
by: Hu, Junhao, et al.
Published: (2025)
S-HPLB: Efficient LLM Attention Serving via Sparsity-Aware Head Parallelism Load Balance
by: Liu, Di, et al.
Published: (2026)
by: Liu, Di, et al.
Published: (2026)
SLO-aware GPU Frequency Scaling for Energy Efficient LLM Inference Serving
by: Kakolyris, Andreas Kosmas, et al.
Published: (2024)
by: Kakolyris, Andreas Kosmas, et al.
Published: (2024)
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
by: Liu, Dong, et al.
Published: (2025)
by: Liu, Dong, et al.
Published: (2025)
MSARS: A Meta-Learning and Reinforcement Learning Framework for SLO Resource Allocation and Adaptive Scaling for Microservices
by: Hu, Kan, et al.
Published: (2024)
by: Hu, Kan, et al.
Published: (2024)
APEX: An Extensible and Dynamism-Aware Simulator for Automated Parallel Execution in LLM Serving
by: Lin, Yi-Chien, et al.
Published: (2024)
by: Lin, Yi-Chien, et al.
Published: (2024)
Accuracy Is Speed: Towards Long-Context-Aware Routing for Distributed LLM Serving
by: Yoshimura, Takeshi, et al.
Published: (2026)
by: Yoshimura, Takeshi, et al.
Published: (2026)
Hummingbird: SLO-Oriented GPU Preemption at Microsecond-scale
by: Hu, Tiancheng, et al.
Published: (2026)
by: Hu, Tiancheng, et al.
Published: (2026)
A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via Faro
by: Jeon, Beomyeol, et al.
Published: (2024)
by: Jeon, Beomyeol, et al.
Published: (2024)
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling
by: Wu, Jingfeng, et al.
Published: (2025)
by: Wu, Jingfeng, et al.
Published: (2025)
An SLO Driven and Cost-Aware Autoscaling Framework for Kubernetes
by: Punniyamoorthy, Vinoth, et al.
Published: (2025)
by: Punniyamoorthy, Vinoth, et al.
Published: (2025)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
by: Chen, Jiabin, et al.
Published: (2024)
by: Chen, Jiabin, et al.
Published: (2024)
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
by: Peng, You, et al.
Published: (2026)
by: Peng, You, et al.
Published: (2026)
CascadeInfer: Length-Aware Scheduling of LLM Serving with Low Latency and Load Balancing
by: Yuan, Yitao, et al.
Published: (2025)
by: Yuan, Yitao, et al.
Published: (2025)
Similar Items
-
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
by: Shen, Haiying, et al.
Published: (2024) -
Training DNN Models over Heterogeneous Clusters with Optimal Performance
by: Nie, Chengyi, et al.
Published: (2024) -
DualScale: Energy-Efficient Disaggregated LLM Serving via Phase-Aware Placement and DVFS
by: Basit, Omar, et al.
Published: (2026) -
PolyServe: Efficient Multi-SLO Serving at Scale
by: Zhu, Kan, et al.
Published: (2025) -
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
by: Hu, Jianmin, et al.
Published: (2025)