gLLM: Global Balanced Pipeline Parallelism System for Distributed LLM Serving with Token Throttling
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Guo, Tianyu, Zhang, Xianwei, Du, Jiangsu, Chen, Zhiguang, Xiao, Nong, Lu, Yutong |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM Inference
par: Zhang, Hongbin, et autres
Publié: (2025)
par: Zhang, Hongbin, et autres
Publié: (2025)
AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System
par: Bai, Fengyao, et autres
Publié: (2026)
par: Bai, Fengyao, et autres
Publié: (2026)
EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration
par: Du, Jiangsu, et autres
Publié: (2025)
par: Du, Jiangsu, et autres
Publié: (2025)
Boosting LLM Serving through Spatial-Temporal GPU Resource Sharing
par: Lin, Zejia, et autres
Publié: (2025)
par: Lin, Zejia, et autres
Publié: (2025)
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
par: Wei, Jinhui, et autres
Publié: (2025)
par: Wei, Jinhui, et autres
Publié: (2025)
PipeMax: Enhancing Offline LLM Inference on Commodity GPU Servers
par: Zhang, Hongbin, et autres
Publié: (2026)
par: Zhang, Hongbin, et autres
Publié: (2026)
S-HPLB: Efficient LLM Attention Serving via Sparsity-Aware Head Parallelism Load Balance
par: Liu, Di, et autres
Publié: (2026)
par: Liu, Di, et autres
Publié: (2026)
RServe: Overlapping Encoding and Prefill for Efficient LMM Inference
par: Guo, Tianyu, et autres
Publié: (2025)
par: Guo, Tianyu, et autres
Publié: (2025)
Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism
par: Srivatsa, Vikranth, et autres
Publié: (2026)
par: Srivatsa, Vikranth, et autres
Publié: (2026)
Tackling the Data-Parallel Load Balancing Bottleneck in LLM Serving: Practical Online Routing at Scale
par: Bu, Tianci, et autres
Publié: (2026)
par: Bu, Tianci, et autres
Publié: (2026)
TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity
par: Lai, Ruiqi, et autres
Publié: (2025)
par: Lai, Ruiqi, et autres
Publié: (2025)
Balancing Pipeline Parallelism with Vocabulary Parallelism
par: Yeung, Man Tsung, et autres
Publié: (2024)
par: Yeung, Man Tsung, et autres
Publié: (2024)
DualMap: Enabling Both Cache Affinity and Load Balancing for Distributed LLM Serving
par: Yuan, Ying, et autres
Publié: (2026)
par: Yuan, Ying, et autres
Publié: (2026)
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
par: Zhou, Qihui, et autres
Publié: (2025)
par: Zhou, Qihui, et autres
Publié: (2025)
Bandwidth-Aware and Cost-Efficient Pipeline Parallel Scheduling in Geo-Distributed LLM Training
par: Zhang, Han, et autres
Publié: (2026)
par: Zhang, Han, et autres
Publié: (2026)
PipeLive: Efficient Live In-place Pipeline Parallelism Reconfiguration for Dynamic LLM Serving
par: Bai, Xu, et autres
Publié: (2026)
par: Bai, Xu, et autres
Publié: (2026)
FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
par: Lin, Yanying, et autres
Publié: (2025)
par: Lin, Yanying, et autres
Publié: (2025)
Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving
par: Cheng, Ke, et autres
Publié: (2024)
par: Cheng, Ke, et autres
Publié: (2024)
APEX: An Extensible and Dynamism-Aware Simulator for Automated Parallel Execution in LLM Serving
par: Lin, Yi-Chien, et autres
Publié: (2024)
par: Lin, Yi-Chien, et autres
Publié: (2024)
Optimizing Long-context LLM Serving via Fine-grained Sequence Parallelism
par: Li, Cong, et autres
Publié: (2025)
par: Li, Cong, et autres
Publié: (2025)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
par: Du, Boxiao, et autres
Publié: (2026)
par: Du, Boxiao, et autres
Publié: (2026)
Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines
par: Wagenländer, Marcel, et autres
Publié: (2026)
par: Wagenländer, Marcel, et autres
Publié: (2026)
Cloud Native System for LLM Inference Serving
par: Xu, Minxian, et autres
Publié: (2025)
par: Xu, Minxian, et autres
Publié: (2025)
WWW.Serve: Interconnecting Global LLM Services through Decentralization
par: Wang, Huanyu, et autres
Publié: (2026)
par: Wang, Huanyu, et autres
Publié: (2026)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
par: He, Yiyuan, et autres
Publié: (2025)
par: He, Yiyuan, et autres
Publié: (2025)
TridentServe: A Stage-level Serving System for Diffusion Pipelines
par: Xia, Yifei, et autres
Publié: (2025)
par: Xia, Yifei, et autres
Publié: (2025)
eLLM: Elastic Memory Management Framework for Efficient LLM Serving
par: Xu, Jiale, et autres
Publié: (2025)
par: Xu, Jiale, et autres
Publié: (2025)
CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration
par: Nian, Sean, et autres
Publié: (2026)
par: Nian, Sean, et autres
Publié: (2026)
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
par: Duan, Jiangfei, et autres
Publié: (2024)
par: Duan, Jiangfei, et autres
Publié: (2024)
CascadeInfer: Length-Aware Scheduling of LLM Serving with Low Latency and Load Balancing
par: Yuan, Yitao, et autres
Publié: (2025)
par: Yuan, Yitao, et autres
Publié: (2025)
MixServe: An Automatic Distributed Serving System for MoE Models with Hybrid Parallelism Based on Fused Communication Algorithm
par: Zhou, Bowen, et autres
Publié: (2026)
par: Zhou, Bowen, et autres
Publié: (2026)
TokenDance: Scaling Multi-Agent LLM Serving via Collective KV Cache Sharing
par: Bian, Zhuohang, et autres
Publié: (2026)
par: Bian, Zhuohang, et autres
Publié: (2026)
OTAS: An Elastic Transformer Serving System via Token Adaptation
par: Chen, Jinyu, et autres
Publié: (2024)
par: Chen, Jinyu, et autres
Publié: (2024)
Argus: Token Aware Distributed LLM Inference Optimization
par: Wu, Panlong, et autres
Publié: (2025)
par: Wu, Panlong, et autres
Publié: (2025)
Jenga: Effective Memory Management for Serving LLM with Heterogeneity
par: Zhang, Chen, et autres
Publié: (2025)
par: Zhang, Chen, et autres
Publié: (2025)
Accuracy Is Speed: Towards Long-Context-Aware Routing for Distributed LLM Serving
par: Yoshimura, Takeshi, et autres
Publié: (2026)
par: Yoshimura, Takeshi, et autres
Publié: (2026)
SiPipe: Bridging the CPU-GPU Utilization Gap for Efficient Pipeline-Parallel LLM Inference
par: He, Yongchao, et autres
Publié: (2025)
par: He, Yongchao, et autres
Publié: (2025)
SPPO:Efficient Long-sequence LLM Training via Adaptive Sequence Pipeline Parallel Offloading
par: Chen, Qiaoling, et autres
Publié: (2025)
par: Chen, Qiaoling, et autres
Publié: (2025)
SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference
par: Xie, Jincheng, et autres
Publié: (2026)
par: Xie, Jincheng, et autres
Publié: (2026)
Predictable LLM Serving on GPU Clusters
par: Darzi, Erfan, et autres
Publié: (2025)
par: Darzi, Erfan, et autres
Publié: (2025)
Documents similaires
-
TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM Inference
par: Zhang, Hongbin, et autres
Publié: (2025) -
AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System
par: Bai, Fengyao, et autres
Publié: (2026) -
EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration
par: Du, Jiangsu, et autres
Publié: (2025) -
Boosting LLM Serving through Spatial-Temporal GPU Resource Sharing
par: Lin, Zejia, et autres
Publié: (2025) -
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
par: Wei, Jinhui, et autres
Publié: (2025)