Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Jiu, Yang, Shuangyan, Xiong, Xu, Duan, Hexiao, Zhang, Xinran, Ren, Jie, Li, Dong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Argus: Token Aware Distributed LLM Inference Optimization
di: Wu, Panlong, et al.
Pubblicazione: (2025)
di: Wu, Panlong, et al.
Pubblicazione: (2025)
Diagonal Scaling: A Multi-Dimensional Resource Model and Optimization Framework for Distributed Databases
di: Abdullah, Shahir, et al.
Pubblicazione: (2025)
di: Abdullah, Shahir, et al.
Pubblicazione: (2025)
Federated Inference for Heterogeneous LLM Communication and Collaboration
di: Chen, Zihan, et al.
Pubblicazione: (2026)
di: Chen, Zihan, et al.
Pubblicazione: (2026)
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
di: Wu, Tian, et al.
Pubblicazione: (2025)
di: Wu, Tian, et al.
Pubblicazione: (2025)
LLM-CoOpt: A Co-Design and Optimization Framework for Efficient LLM Inference on Heterogeneous Platforms
di: Kong, Jie, et al.
Pubblicazione: (2026)
di: Kong, Jie, et al.
Pubblicazione: (2026)
CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism
di: Ma, Bin, et al.
Pubblicazione: (2026)
di: Ma, Bin, et al.
Pubblicazione: (2026)
OnePiece: A Large-Scale Distributed Inference System with RDMA for Complex AI-Generated Content (AIGC) Workflows
di: Chen, June, et al.
Pubblicazione: (2026)
di: Chen, June, et al.
Pubblicazione: (2026)
Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
di: Wu, Yu, et al.
Pubblicazione: (2025)
di: Wu, Yu, et al.
Pubblicazione: (2025)
Distributed Inference Performance Optimization for LLMs on CPUs
di: He, Pujiang, et al.
Pubblicazione: (2024)
di: He, Pujiang, et al.
Pubblicazione: (2024)
Characterizing Communication Patterns in Distributed Large Language Model Inference
di: Xu, Lang, et al.
Pubblicazione: (2025)
di: Xu, Lang, et al.
Pubblicazione: (2025)
AnchorTP: Resilient LLM Inference with State-Preserving Elastic Tensor Parallelism
di: Xu, Wendong, et al.
Pubblicazione: (2025)
di: Xu, Wendong, et al.
Pubblicazione: (2025)
Scaling LLM Inference Beyond Amdahl`s Limits via Eliminating Non-Scalable Overheads
di: Zhao, Alan, et al.
Pubblicazione: (2026)
di: Zhao, Alan, et al.
Pubblicazione: (2026)
Lagom: Unleashing the Power of Communication and Computation Overlapping for Distributed LLM Training
di: Xu, Guanbin, et al.
Pubblicazione: (2026)
di: Xu, Guanbin, et al.
Pubblicazione: (2026)
Distributed On-Device LLM Inference With Over-the-Air Computation
di: Zhang, Kai, et al.
Pubblicazione: (2025)
di: Zhang, Kai, et al.
Pubblicazione: (2025)
Efficient Multi-round LLM Inference over Disaggregated Serving
di: He, Wenhao, et al.
Pubblicazione: (2026)
di: He, Wenhao, et al.
Pubblicazione: (2026)
Exploring and Evaluating Real-world CXL: Use Cases and System Adoption
di: Wang, Xi, et al.
Pubblicazione: (2024)
di: Wang, Xi, et al.
Pubblicazione: (2024)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
Communication-Efficient Collaborative LLM Inference over LEO Satellite Networks
di: Zhang, Songge, et al.
Pubblicazione: (2026)
di: Zhang, Songge, et al.
Pubblicazione: (2026)
VQ-LLM: High-performance Code Generation for Vector Quantization Augmented LLM Inference
di: Liu, Zihan, et al.
Pubblicazione: (2025)
di: Liu, Zihan, et al.
Pubblicazione: (2025)
ReviveMoE: Fast Recovery for Hardware Failures in Large-Scale MoE LLM Inference Deployments
di: Li, Haley, et al.
Pubblicazione: (2026)
di: Li, Haley, et al.
Pubblicazione: (2026)
AMSP: Reducing Communication Overhead of ZeRO for Efficient LLM Training
di: Chen, Qiaoling, et al.
Pubblicazione: (2023)
di: Chen, Qiaoling, et al.
Pubblicazione: (2023)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
di: Xu, Chuhao, et al.
Pubblicazione: (2025)
di: Xu, Chuhao, et al.
Pubblicazione: (2025)
Hyperion: Hierarchical Scheduling for Parallel LLM Acceleration in Multi-tier Networks
di: Ma, Mulei, et al.
Pubblicazione: (2025)
di: Ma, Mulei, et al.
Pubblicazione: (2025)
DWDP: Distributed Weight Data Parallelism for High-Performance LLM Inference on NVL72
di: Li, Wanqian, et al.
Pubblicazione: (2026)
di: Li, Wanqian, et al.
Pubblicazione: (2026)
Cloud Native System for LLM Inference Serving
di: Xu, Minxian, et al.
Pubblicazione: (2025)
di: Xu, Minxian, et al.
Pubblicazione: (2025)
λScale: Enabling Fast Scaling for Serverless Large Language Model Inference
di: Yu, Minchen, et al.
Pubblicazione: (2025)
di: Yu, Minchen, et al.
Pubblicazione: (2025)
Optimizing Distributed Training Approaches for Scaling Neural Networks
di: Baligodugula, Vishnu Vardhan, et al.
Pubblicazione: (2025)
di: Baligodugula, Vishnu Vardhan, et al.
Pubblicazione: (2025)
MegatronApp: Efficient and Comprehensive Management on Distributed LLM Training
di: Zhao, Bohan, et al.
Pubblicazione: (2025)
di: Zhao, Bohan, et al.
Pubblicazione: (2025)
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference
di: Zhao, Yihao, et al.
Pubblicazione: (2025)
di: Zhao, Yihao, et al.
Pubblicazione: (2025)
TokenDance: Scaling Multi-Agent LLM Serving via Collective KV Cache Sharing
di: Bian, Zhuohang, et al.
Pubblicazione: (2026)
di: Bian, Zhuohang, et al.
Pubblicazione: (2026)
Communication-Efficient Distributed On-Device LLM Inference Over Wireless Networks
di: Zhang, Kai, et al.
Pubblicazione: (2025)
di: Zhang, Kai, et al.
Pubblicazione: (2025)
TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference
di: Gond, Raja, et al.
Pubblicazione: (2025)
di: Gond, Raja, et al.
Pubblicazione: (2025)
DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance
di: Zhang, Yuning, et al.
Pubblicazione: (2025)
di: Zhang, Yuning, et al.
Pubblicazione: (2025)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
di: Karfakis, George, et al.
Pubblicazione: (2025)
di: Karfakis, George, et al.
Pubblicazione: (2025)
MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production
di: Xue, Chunyu, et al.
Pubblicazione: (2026)
di: Xue, Chunyu, et al.
Pubblicazione: (2026)
On-the-fly Communication-and-Computing to Enable Representation Learning for Distributed Point Clouds
di: Chen, Xu, et al.
Pubblicazione: (2024)
di: Chen, Xu, et al.
Pubblicazione: (2024)
ACE-Sync: An Adaptive Cloud-Edge Synchronization Framework for Communication-Efficient Large-Scale Distributed Model Training
di: Yang, Yi, et al.
Pubblicazione: (2025)
di: Yang, Yi, et al.
Pubblicazione: (2025)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
di: He, Yiyuan, et al.
Pubblicazione: (2024)
di: He, Yiyuan, et al.
Pubblicazione: (2024)
Understanding and Improving Communication Performance in Multi-node LLM Inference
di: Singhania, Prajwal, et al.
Pubblicazione: (2025)
di: Singhania, Prajwal, et al.
Pubblicazione: (2025)
MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services
di: Yu, Dianhai, et al.
Pubblicazione: (2022)
di: Yu, Dianhai, et al.
Pubblicazione: (2022)
Documenti analoghi
-
Argus: Token Aware Distributed LLM Inference Optimization
di: Wu, Panlong, et al.
Pubblicazione: (2025) -
Diagonal Scaling: A Multi-Dimensional Resource Model and Optimization Framework for Distributed Databases
di: Abdullah, Shahir, et al.
Pubblicazione: (2025) -
Federated Inference for Heterogeneous LLM Communication and Collaboration
di: Chen, Zihan, et al.
Pubblicazione: (2026) -
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
di: Wu, Tian, et al.
Pubblicazione: (2025) -
LLM-CoOpt: A Co-Design and Optimization Framework for Efficient LLM Inference on Heterogeneous Platforms
di: Kong, Jie, et al.
Pubblicazione: (2026)