Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
Fuente:
arXiv
Saved in:
| Main Author: | Wang, Qipeng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SLO-Aware Scheduling for Large Language Model Inferences
by: Huang, Jinqi, et al.
Published: (2025)
by: Huang, Jinqi, et al.
Published: (2025)
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
by: Chow, Will
Published: (2025)
by: Chow, Will
Published: (2025)
Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
by: Wu, Yu, et al.
Published: (2025)
by: Wu, Yu, et al.
Published: (2025)
SLO-Aware Compute Resource Allocation for Prefill-Decode Disaggregated LLM Inference
by: Li, Luchang, et al.
Published: (2026)
by: Li, Luchang, et al.
Published: (2026)
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
by: Li, Rongzhi, et al.
Published: (2025)
by: Li, Rongzhi, et al.
Published: (2025)
SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
by: Cheng, Ke, et al.
Published: (2024)
by: Cheng, Ke, et al.
Published: (2024)
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
by: Nie, Chengyi, et al.
Published: (2024)
by: Nie, Chengyi, et al.
Published: (2024)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
by: Hu, Cunchen, et al.
Published: (2024)
by: Hu, Cunchen, et al.
Published: (2024)
PromptTuner: SLO-Aware Elastic System for LLM Prompt Tuning
by: Gao, Wei, et al.
Published: (2026)
by: Gao, Wei, et al.
Published: (2026)
SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips
by: Yu, Jiahuan, et al.
Published: (2026)
by: Yu, Jiahuan, et al.
Published: (2026)
From Attention to Disaggregation: Tracing the Evolution of LLM Inference
by: Kumar, Madabattula Rajesh, et al.
Published: (2025)
by: Kumar, Madabattula Rajesh, et al.
Published: (2025)
PROSERVE: Unified Multi-Priority Request Scheduling for LLM Serving
by: Huang, Weizhe, et al.
Published: (2025)
by: Huang, Weizhe, et al.
Published: (2025)
Efficient Multi-round LLM Inference over Disaggregated Serving
by: He, Wenhao, et al.
Published: (2026)
by: He, Wenhao, et al.
Published: (2026)
CASA: A Framework for SLO and Carbon-Aware Autoscaling and Scheduling in Serverless Cloud Computing
by: Qi, S., et al.
Published: (2024)
by: Qi, S., et al.
Published: (2024)
SLO-Aware Task Offloading within Collaborative Vehicle Platoons
by: Sedlak, Boris, et al.
Published: (2024)
by: Sedlak, Boris, et al.
Published: (2024)
A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via Faro
by: Jeon, Beomyeol, et al.
Published: (2024)
by: Jeon, Beomyeol, et al.
Published: (2024)
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
by: Ma, Chenxiang, et al.
Published: (2025)
by: Ma, Chenxiang, et al.
Published: (2025)
JITServe: SLO-aware LLM Serving with Imprecise Request Information
by: Zhang, Wei, et al.
Published: (2025)
by: Zhang, Wei, et al.
Published: (2025)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
by: Hu, Jianmin, et al.
Published: (2025)
by: Hu, Jianmin, et al.
Published: (2025)
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
by: Liao, Junhan, et al.
Published: (2025)
by: Liao, Junhan, et al.
Published: (2025)
TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM Inference
by: Zhang, Hongbin, et al.
Published: (2025)
by: Zhang, Hongbin, et al.
Published: (2025)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
by: Chen, Jiabin, et al.
Published: (2024)
by: Chen, Jiabin, et al.
Published: (2024)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
by: Pan, Xinglin, et al.
Published: (2025)
by: Pan, Xinglin, et al.
Published: (2025)
Efficient Routing of Inference Requests across LLM Instances in Cloud-Edge Computing
by: Yu, Shibo, et al.
Published: (2025)
by: Yu, Shibo, et al.
Published: (2025)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
by: Zhang, Zhexiang, et al.
Published: (2025)
by: Zhang, Zhexiang, et al.
Published: (2025)
Harli: SLO-Aware Co-location of LLM Inference and PEFT-based Finetuning on Model-as-a-Service Platforms
by: Xu, Ao, et al.
Published: (2025)
by: Xu, Ao, et al.
Published: (2025)
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
by: Shen, Haiying, et al.
Published: (2024)
by: Shen, Haiying, et al.
Published: (2024)
Lotus: Optimizing Disaggregated Transactions with Disaggregated Locks
by: Hu, Zhisheng, et al.
Published: (2025)
by: Hu, Zhisheng, et al.
Published: (2025)
SIMPLE: Disaggregating Sampling from GPU Inference into a Decision Plane for Faster Distributed LLM Serving
by: Zhao, Bohan, et al.
Published: (2025)
by: Zhao, Bohan, et al.
Published: (2025)
DualScale: Energy-Efficient Disaggregated LLM Serving via Phase-Aware Placement and DVFS
by: Basit, Omar, et al.
Published: (2026)
by: Basit, Omar, et al.
Published: (2026)
HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving
by: Dong, Xianzhe, et al.
Published: (2025)
by: Dong, Xianzhe, et al.
Published: (2025)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
by: Gu, Jianfeng, et al.
Published: (2025)
by: Gu, Jianfeng, et al.
Published: (2025)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
by: Mo, Zizhao, et al.
Published: (2026)
by: Mo, Zizhao, et al.
Published: (2026)
KVDirect: Distributed Disaggregated LLM Inference
by: Chen, Shiyang, et al.
Published: (2024)
by: Chen, Shiyang, et al.
Published: (2024)
Hummingbird: SLO-Oriented GPU Preemption at Microsecond-scale
by: Hu, Tiancheng, et al.
Published: (2026)
by: Hu, Tiancheng, et al.
Published: (2026)
HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment
by: Jiang, Youhe, et al.
Published: (2025)
by: Jiang, Youhe, et al.
Published: (2025)
SERFLOW: A Cross-Service Cost Optimization Framework for SLO-Aware Dynamic ML Inference
by: Zhang, Zongshun, et al.
Published: (2025)
by: Zhang, Zongshun, et al.
Published: (2025)
PARD: Enhancing Goodput for Inference Pipeline via Proactive Request Dropping
by: Zhao, Zhixin, et al.
Published: (2026)
by: Zhao, Zhixin, et al.
Published: (2026)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
by: Hu, Cunchen, et al.
Published: (2024)
by: Hu, Cunchen, et al.
Published: (2024)
Argus: Token Aware Distributed LLM Inference Optimization
by: Wu, Panlong, et al.
Published: (2025)
by: Wu, Panlong, et al.
Published: (2025)
Similar Items
-
SLO-Aware Scheduling for Large Language Model Inferences
by: Huang, Jinqi, et al.
Published: (2025) -
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
by: Chow, Will
Published: (2025) -
Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
by: Wu, Yu, et al.
Published: (2025) -
SLO-Aware Compute Resource Allocation for Prefill-Decode Disaggregated LLM Inference
by: Li, Luchang, et al.
Published: (2026) -
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
by: Li, Rongzhi, et al.
Published: (2025)