Scaling LLM Inference Beyond Amdahl`s Limits via Eliminating Non-Scalable Overheads
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhao, Alan, He, Cyril Y., Xu, Wei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SiDP: Memory-Efficient Data Parallelism for Offline LLM Inference
por: Zhao, Alan, et al.
Publicado: (2026)
por: Zhao, Alan, et al.
Publicado: (2026)
Amdahl's and Gustafson-Barsis laws revisited
por: Karbowski, Andrzej
Publicado: (2008)
por: Karbowski, Andrzej
Publicado: (2008)
AMSP: Reducing Communication Overhead of ZeRO for Efficient LLM Training
por: Chen, Qiaoling, et al.
Publicado: (2023)
por: Chen, Qiaoling, et al.
Publicado: (2023)
Modernizing Amdahl's Law: How AI Scaling Laws Shape Computer Architecture
por: Lu, Chien-Ping
Publicado: (2026)
por: Lu, Chien-Ping
Publicado: (2026)
Cloud Native System for LLM Inference Serving
por: Xu, Minxian, et al.
Publicado: (2025)
por: Xu, Minxian, et al.
Publicado: (2025)
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling
por: Wu, Jingfeng, et al.
Publicado: (2025)
por: Wu, Jingfeng, et al.
Publicado: (2025)
LLM-Emu: Native Runtime Emulation of LLM Inference via Profile-Driven Sampling
por: Da, Wei, et al.
Publicado: (2026)
por: Da, Wei, et al.
Publicado: (2026)
Efficient Multi-round LLM Inference over Disaggregated Serving
por: He, Wenhao, et al.
Publicado: (2026)
por: He, Wenhao, et al.
Publicado: (2026)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
por: He, Yiyuan, et al.
Publicado: (2024)
por: He, Yiyuan, et al.
Publicado: (2024)
Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization
por: Chen, Jiu, et al.
Publicado: (2026)
por: Chen, Jiu, et al.
Publicado: (2026)
Checkmate: Zero-Overhead Model Checkpointing via Network Gradient Replication
por: Bhardwaj, Ankit, et al.
Publicado: (2025)
por: Bhardwaj, Ankit, et al.
Publicado: (2025)
SiPipe: Bridging the CPU-GPU Utilization Gap for Efficient Pipeline-Parallel LLM Inference
por: He, Yongchao, et al.
Publicado: (2025)
por: He, Yongchao, et al.
Publicado: (2025)
Collaborative Inference Acceleration with Non-Penetrative Tensor Partitioning
por: Liu, Zhibang, et al.
Publicado: (2025)
por: Liu, Zhibang, et al.
Publicado: (2025)
SIMPLE: Disaggregating Sampling from GPU Inference into a Decision Plane for Faster Distributed LLM Serving
por: Zhao, Bohan, et al.
Publicado: (2025)
por: Zhao, Bohan, et al.
Publicado: (2025)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
por: Xu, Chuhao, et al.
Publicado: (2025)
por: Xu, Chuhao, et al.
Publicado: (2025)
RcLLM: Accelerating Generative Recommendation via Beyond-Prefix KV Caching
por: Zhao, Zhan, et al.
Publicado: (2026)
por: Zhao, Zhan, et al.
Publicado: (2026)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
por: Hu, Cunchen, et al.
Publicado: (2024)
por: Hu, Cunchen, et al.
Publicado: (2024)
Near-Zero-Overhead Freshness for Recommendation Systems via Inference-Side Model Updates
por: Yu, Wenjun, et al.
Publicado: (2025)
por: Yu, Wenjun, et al.
Publicado: (2025)
DualScale: Energy-Efficient Disaggregated LLM Serving via Phase-Aware Placement and DVFS
por: Basit, Omar, et al.
Publicado: (2026)
por: Basit, Omar, et al.
Publicado: (2026)
TaxBreak: Unmasking the Hidden Costs of LLM Inference Through Overhead Decomposition
por: Vellaisamy, Prabhu, et al.
Publicado: (2026)
por: Vellaisamy, Prabhu, et al.
Publicado: (2026)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
por: Zhang, Zhexiang, et al.
Publicado: (2025)
por: Zhang, Zhexiang, et al.
Publicado: (2025)
Distributed On-Device LLM Inference With Over-the-Air Computation
por: Zhang, Kai, et al.
Publicado: (2025)
por: Zhang, Kai, et al.
Publicado: (2025)
Beyond Microservices: Testing Web-Scale RCA Methods on GPU-Driven LLM Workloads
por: Scheinert, Dominik, et al.
Publicado: (2026)
por: Scheinert, Dominik, et al.
Publicado: (2026)
Modular Architecture for High-Performance and Low Overhead Data Transfers
por: Swargo, Rasman Mubtasim, et al.
Publicado: (2025)
por: Swargo, Rasman Mubtasim, et al.
Publicado: (2025)
Cronus: Efficient LLM inference on Heterogeneous GPU Clusters via Partially Disaggregated Prefill
por: Liu, Yunzhao, et al.
Publicado: (2025)
por: Liu, Yunzhao, et al.
Publicado: (2025)
KaMPIng: Flexible and (Near) Zero-Overhead C++ Bindings for MPI
por: Uhl, Tim Niklas, et al.
Publicado: (2024)
por: Uhl, Tim Niklas, et al.
Publicado: (2024)
EdgeShard: Efficient LLM Inference via Collaborative Edge Computing
por: Zhang, Mingjin, et al.
Publicado: (2024)
por: Zhang, Mingjin, et al.
Publicado: (2024)
Quantifying the Energy Consumption and Carbon Emissions of LLM Inference via Simulations
por: Özcan, Miray, et al.
Publicado: (2025)
por: Özcan, Miray, et al.
Publicado: (2025)
LLM-CoOpt: A Co-Design and Optimization Framework for Efficient LLM Inference on Heterogeneous Platforms
por: Kong, Jie, et al.
Publicado: (2026)
por: Kong, Jie, et al.
Publicado: (2026)
AnchorTP: Resilient LLM Inference with State-Preserving Elastic Tensor Parallelism
por: Xu, Wendong, et al.
Publicado: (2025)
por: Xu, Wendong, et al.
Publicado: (2025)
λScale: Enabling Fast Scaling for Serverless Large Language Model Inference
por: Yu, Minchen, et al.
Publicado: (2025)
por: Yu, Minchen, et al.
Publicado: (2025)
MegatronApp: Efficient and Comprehensive Management on Distributed LLM Training
por: Zhao, Bohan, et al.
Publicado: (2025)
por: Zhao, Bohan, et al.
Publicado: (2025)
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
por: Chen, Haoyu, et al.
Publicado: (2025)
por: Chen, Haoyu, et al.
Publicado: (2025)
Ripple: Scalable Incremental GNN Inferencing on Large Streaming Graphs
por: Naman, Pranjal, et al.
Publicado: (2025)
por: Naman, Pranjal, et al.
Publicado: (2025)
Understanding and Reducing Metadata-Driven Host Overheads in Sampling-Based GNN Training
por: Gong, Yidong, et al.
Publicado: (2026)
por: Gong, Yidong, et al.
Publicado: (2026)
Parallelize Over Data Particle Advection: Participation, Ping Pong Particles, and Overhead
por: Wang, Zhe, et al.
Publicado: (2024)
por: Wang, Zhe, et al.
Publicado: (2024)
SDSL-Solver: Scalable Distributed Sparse Linear Solvers for Large-Scale Interior Point Methods
por: Yang, Shaofeng, et al.
Publicado: (2026)
por: Yang, Shaofeng, et al.
Publicado: (2026)
Shift Parallelism: Low-Latency, High-Throughput LLM Inference for Dynamic Workloads
por: Hidayetoglu, Mert, et al.
Publicado: (2025)
por: Hidayetoglu, Mert, et al.
Publicado: (2025)
ReviveMoE: Fast Recovery for Hardware Failures in Large-Scale MoE LLM Inference Deployments
por: Li, Haley, et al.
Publicado: (2026)
por: Li, Haley, et al.
Publicado: (2026)
Scaling Up Throughput-oriented LLM Inference Applications on Heterogeneous Opportunistic GPU Clusters with Pervasive Context Management
por: Phung, Thanh Son, et al.
Publicado: (2025)
por: Phung, Thanh Son, et al.
Publicado: (2025)
Ejemplares similares
-
SiDP: Memory-Efficient Data Parallelism for Offline LLM Inference
por: Zhao, Alan, et al.
Publicado: (2026) -
Amdahl's and Gustafson-Barsis laws revisited
por: Karbowski, Andrzej
Publicado: (2008) -
AMSP: Reducing Communication Overhead of ZeRO for Efficient LLM Training
por: Chen, Qiaoling, et al.
Publicado: (2023) -
Modernizing Amdahl's Law: How AI Scaling Laws Shape Computer Architecture
por: Lu, Chien-Ping
Publicado: (2026) -
Cloud Native System for LLM Inference Serving
por: Xu, Minxian, et al.
Publicado: (2025)