A Universal Load Balancing Principle and Its Application to Large Language Model Serving
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Zixi, Bu, Tianci, Song, Chendong, Lu, Xin, Ye, Yinyu, Zhou, Zijie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Tackling the Data-Parallel Load Balancing Bottleneck in LLM Serving: Practical Online Routing at Scale
von: Bu, Tianci, et al.
Veröffentlicht: (2026)
von: Bu, Tianci, et al.
Veröffentlicht: (2026)
DualMap: Enabling Both Cache Affinity and Load Balancing for Distributed LLM Serving
von: Yuan, Ying, et al.
Veröffentlicht: (2026)
von: Yuan, Ying, et al.
Veröffentlicht: (2026)
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
von: Zhong, Yinmin, et al.
Veröffentlicht: (2024)
von: Zhong, Yinmin, et al.
Veröffentlicht: (2024)
ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production
von: Xiang, Yuxing, et al.
Veröffentlicht: (2025)
von: Xiang, Yuxing, et al.
Veröffentlicht: (2025)
Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving
von: Cheng, Ke, et al.
Veröffentlicht: (2024)
von: Cheng, Ke, et al.
Veröffentlicht: (2024)
S-HPLB: Efficient LLM Attention Serving via Sparsity-Aware Head Parallelism Load Balance
von: Liu, Di, et al.
Veröffentlicht: (2026)
von: Liu, Di, et al.
Veröffentlicht: (2026)
DeepServe: Serverless Large Language Model Serving at Scale
von: Hu, Junhao, et al.
Veröffentlicht: (2025)
von: Hu, Junhao, et al.
Veröffentlicht: (2025)
CascadeInfer: Length-Aware Scheduling of LLM Serving with Low Latency and Load Balancing
von: Yuan, Yitao, et al.
Veröffentlicht: (2025)
von: Yuan, Yitao, et al.
Veröffentlicht: (2025)
Position: LLM Serving Needs Mathematical Optimization and Algorithmic Foundations, Not Just Heuristics
von: Zhou, Zijie
Veröffentlicht: (2026)
von: Zhou, Zijie
Veröffentlicht: (2026)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
von: Chen, Xing, et al.
Veröffentlicht: (2025)
von: Chen, Xing, et al.
Veröffentlicht: (2025)
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
von: Qianli, Liu, et al.
Veröffentlicht: (2025)
von: Qianli, Liu, et al.
Veröffentlicht: (2025)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
von: He, Yiyuan, et al.
Veröffentlicht: (2025)
von: He, Yiyuan, et al.
Veröffentlicht: (2025)
NanoFlow: Towards Optimal Large Language Model Serving Throughput
von: Zhu, Kan, et al.
Veröffentlicht: (2024)
von: Zhu, Kan, et al.
Veröffentlicht: (2024)
Cascadia: An Efficient Cascade Serving System for Large Language Models
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
Serving Chain-structured Jobs with Large Memory Footprints with Application to Large Foundation Model Serving
von: Sun, Tingyang, et al.
Veröffentlicht: (2026)
von: Sun, Tingyang, et al.
Veröffentlicht: (2026)
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
von: Chen, Hongyu, et al.
Veröffentlicht: (2026)
von: Chen, Hongyu, et al.
Veröffentlicht: (2026)
gLLM: Global Balanced Pipeline Parallelism System for Distributed LLM Serving with Token Throttling
von: Guo, Tianyu, et al.
Veröffentlicht: (2025)
von: Guo, Tianyu, et al.
Veröffentlicht: (2025)
FLYING SERVING: On-the-Fly Parallelism Switching for Large Language Model Serving
von: Gao, Shouwei, et al.
Veröffentlicht: (2026)
von: Gao, Shouwei, et al.
Veröffentlicht: (2026)
Block: Balancing Load in LLM Serving with Context, Knowledge and Predictive Scheduling
von: Da, Wei, et al.
Veröffentlicht: (2025)
von: Da, Wei, et al.
Veröffentlicht: (2025)
Towards Sustainable Large Language Model Serving
von: Nguyen, Sophia, et al.
Veröffentlicht: (2024)
von: Nguyen, Sophia, et al.
Veröffentlicht: (2024)
HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving
von: Dong, Xianzhe, et al.
Veröffentlicht: (2025)
von: Dong, Xianzhe, et al.
Veröffentlicht: (2025)
BlockLLM: Multi-tenant Finer-grained Serving for Large Language Models
von: Hu, Bodun, et al.
Veröffentlicht: (2024)
von: Hu, Bodun, et al.
Veröffentlicht: (2024)
Fine-grained MoE Load Balancing with Linear Programming
von: Zhao, Chenqi, et al.
Veröffentlicht: (2025)
von: Zhao, Chenqi, et al.
Veröffentlicht: (2025)
Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities
von: Chen, Zhixiong, et al.
Veröffentlicht: (2026)
von: Chen, Zhixiong, et al.
Veröffentlicht: (2026)
TD-Orch: Scalable Load-Balancing for Distributed Systems with Applications to Graph Processing
von: Zhao, Yiwei, et al.
Veröffentlicht: (2025)
von: Zhao, Yiwei, et al.
Veröffentlicht: (2025)
LB4OMP: A Dynamic Load Balancing Library for Multithreaded Applications
von: Korndörfer, Jonas H. Müller, et al.
Veröffentlicht: (2021)
von: Korndörfer, Jonas H. Müller, et al.
Veröffentlicht: (2021)
Kairos: Low-latency Multi-Agent Serving with Shared LLMs and Excessive Loads in the Public Cloud
von: Chen, Jinyuan, et al.
Veröffentlicht: (2025)
von: Chen, Jinyuan, et al.
Veröffentlicht: (2025)
LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism
von: Wu, Bingyang, et al.
Veröffentlicht: (2024)
von: Wu, Bingyang, et al.
Veröffentlicht: (2024)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
von: Mo, Zizhao, et al.
Veröffentlicht: (2026)
von: Mo, Zizhao, et al.
Veröffentlicht: (2026)
PlanetServe: A Decentralized, Scalable, and Privacy-Preserving Overlay for Democratizing Large Language Model Serving
von: Fang, Fei, et al.
Veröffentlicht: (2025)
von: Fang, Fei, et al.
Veröffentlicht: (2025)
MoLink: Distributed and Efficient Serving Framework for Large Models
von: Jin, Lewei, et al.
Veröffentlicht: (2025)
von: Jin, Lewei, et al.
Veröffentlicht: (2025)
Fast Distributed Inference Serving for Large Language Models
von: Wu, Bingyang, et al.
Veröffentlicht: (2023)
von: Wu, Bingyang, et al.
Veröffentlicht: (2023)
A Communication- and Memory-Aware Model for Load Balancing Tasks
von: Lifflander, Jonathan, et al.
Veröffentlicht: (2024)
von: Lifflander, Jonathan, et al.
Veröffentlicht: (2024)
NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor Parallelism
von: Ai, Xin, et al.
Veröffentlicht: (2024)
von: Ai, Xin, et al.
Veröffentlicht: (2024)
GENSERVE: Efficient Co-Serving of Heterogeneous Diffusion Model Workloads
von: Ye, Fanjiang, et al.
Veröffentlicht: (2026)
von: Ye, Fanjiang, et al.
Veröffentlicht: (2026)
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models
von: Wang, Wei, et al.
Veröffentlicht: (2024)
von: Wang, Wei, et al.
Veröffentlicht: (2024)
Equinox: Holistic Fair Scheduling in Serving Large Language Models
von: Wei, Zhixiang, et al.
Veröffentlicht: (2025)
von: Wei, Zhixiang, et al.
Veröffentlicht: (2025)
AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System
von: Bai, Fengyao, et al.
Veröffentlicht: (2026)
von: Bai, Fengyao, et al.
Veröffentlicht: (2026)
Stateful Large Language Model Serving with Pensieve
von: Yu, Lingfan, et al.
Veröffentlicht: (2023)
von: Yu, Lingfan, et al.
Veröffentlicht: (2023)
EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration
von: Du, Jiangsu, et al.
Veröffentlicht: (2025)
von: Du, Jiangsu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Tackling the Data-Parallel Load Balancing Bottleneck in LLM Serving: Practical Online Routing at Scale
von: Bu, Tianci, et al.
Veröffentlicht: (2026) -
DualMap: Enabling Both Cache Affinity and Load Balancing for Distributed LLM Serving
von: Yuan, Ying, et al.
Veröffentlicht: (2026) -
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
von: Zhong, Yinmin, et al.
Veröffentlicht: (2024) -
ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production
von: Xiang, Yuxing, et al.
Veröffentlicht: (2025) -
Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving
von: Cheng, Ke, et al.
Veröffentlicht: (2024)