Distributed Inference Performance Optimization for LLMs on CPUs
Fuente:
arXiv
Saved in:
| Main Authors: | He, Pujiang, Zhou, Shan, Li, Changqing, Huang, Wenhuan, Yu, Weifei, Wang, Duyi, Meng, Chen, Gui, Sheng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Inference Performance Optimization for Large Language Models on CPUs
by: He, Pujiang, et al.
Published: (2024)
by: He, Pujiang, et al.
Published: (2024)
Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference
by: Zhang, Haolin, et al.
Published: (2025)
by: Zhang, Haolin, et al.
Published: (2025)
Inference Acceleration for Large Language Models on CPUs
by: PS, Ditto, et al.
Published: (2024)
by: PS, Ditto, et al.
Published: (2024)
Analyzing the Performance Portability of SYCL across CPUs, GPUs, and Hybrid Systems with SW Sequence Alignment
by: Costanzo, Manuel, et al.
Published: (2024)
by: Costanzo, Manuel, et al.
Published: (2024)
Argus: Token Aware Distributed LLM Inference Optimization
by: Wu, Panlong, et al.
Published: (2025)
by: Wu, Panlong, et al.
Published: (2025)
Towards High-Performance and Portable Molecular Docking on CPUs through Vectorization
by: Accordi, Gianmarco, et al.
Published: (2025)
by: Accordi, Gianmarco, et al.
Published: (2025)
Joint Optimization of Offloading, Batching and DVFS for Multiuser Co-Inference
by: Xu, Yaodan, et al.
Published: (2025)
by: Xu, Yaodan, et al.
Published: (2025)
ArcLight: A Lightweight LLM Inference Architecture for Many-Core CPUs
by: Xu, Yuzhuang, et al.
Published: (2026)
by: Xu, Yuzhuang, et al.
Published: (2026)
A dynamic parallel method for performance optimization on hybrid CPUs
by: Yu, Luo, et al.
Published: (2024)
by: Yu, Luo, et al.
Published: (2024)
SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
by: Cheng, Ke, et al.
Published: (2024)
by: Cheng, Ke, et al.
Published: (2024)
LLM-CoOpt: A Co-Design and Optimization Framework for Efficient LLM Inference on Heterogeneous Platforms
by: Kong, Jie, et al.
Published: (2026)
by: Kong, Jie, et al.
Published: (2026)
MAGNUS: Generating Data Locality to Accelerate Sparse Matrix-Matrix Multiplication on CPUs
by: Wolfson-Pou, Jordi, et al.
Published: (2025)
by: Wolfson-Pou, Jordi, et al.
Published: (2025)
Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization
by: Chen, Jiu, et al.
Published: (2026)
by: Chen, Jiu, et al.
Published: (2026)
ScalableHD: Scalable and High-Throughput Hyperdimensional Computing Inference on Multi-Core CPUs
by: Parikh, Dhruv, et al.
Published: (2025)
by: Parikh, Dhruv, et al.
Published: (2025)
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
by: Wu, Tian, et al.
Published: (2025)
by: Wu, Tian, et al.
Published: (2025)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
by: Hu, Cunchen, et al.
Published: (2024)
by: Hu, Cunchen, et al.
Published: (2024)
OnePiece: A Large-Scale Distributed Inference System with RDMA for Complex AI-Generated Content (AIGC) Workflows
by: Chen, June, et al.
Published: (2026)
by: Chen, June, et al.
Published: (2026)
Online Optimization of DNN Inference Network Utility in Collaborative Edge Computing
by: Li, Rui, et al.
Published: (2024)
by: Li, Rui, et al.
Published: (2024)
GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference
by: Tran, Phuong, et al.
Published: (2025)
by: Tran, Phuong, et al.
Published: (2025)
Comparison of Vectorization Capabilities of Different Compilers for X86 and ARM CPUs
by: Sakib, Nazmus, et al.
Published: (2025)
by: Sakib, Nazmus, et al.
Published: (2025)
Distributed Locking: Performance Analysis and Optimization Strategies
by: Rodriguez, Andre, et al.
Published: (2025)
by: Rodriguez, Andre, et al.
Published: (2025)
SLO-Aware Scheduling for Large Language Model Inferences
by: Huang, Jinqi, et al.
Published: (2025)
by: Huang, Jinqi, et al.
Published: (2025)
FUSCO: High-Performance Distributed Data Shuffling via Transformation-Communication Fusion
by: Zhu, Zhuoran, et al.
Published: (2025)
by: Zhu, Zhuoran, et al.
Published: (2025)
FlexPie: Accelerate Distributed Inference on Edge Devices with Flexible Combinatorial Optimization[Technical Report]
by: Zhang, Runhua, et al.
Published: (2025)
by: Zhang, Runhua, et al.
Published: (2025)
Accelerating Distributed MoE Training and Inference with Lina
by: Li, Jiamin, et al.
Published: (2022)
by: Li, Jiamin, et al.
Published: (2022)
Distributed On-Device LLM Inference With Over-the-Air Computation
by: Zhang, Kai, et al.
Published: (2025)
by: Zhang, Kai, et al.
Published: (2025)
Hybrid-Parallel: Achieving High Performance and Energy Efficient Distributed Inference on Robots
by: Sun, Zekai, et al.
Published: (2024)
by: Sun, Zekai, et al.
Published: (2024)
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
by: Huang, En-Ming, et al.
Published: (2025)
by: Huang, En-Ming, et al.
Published: (2025)
Are Bus-Mounted Edge Servers Feasible?
by: Li, Xuezhi, et al.
Published: (2025)
by: Li, Xuezhi, et al.
Published: (2025)
KV Cache Compression for Inference Efficiency in LLMs: A Review
by: Liu, Yanyu, et al.
Published: (2025)
by: Liu, Yanyu, et al.
Published: (2025)
CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism
by: Ma, Bin, et al.
Published: (2026)
by: Ma, Bin, et al.
Published: (2026)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
by: Li, Suyi, et al.
Published: (2024)
by: Li, Suyi, et al.
Published: (2024)
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
by: Wang, Zhibin, et al.
Published: (2025)
by: Wang, Zhibin, et al.
Published: (2025)
CoCoI: Distributed Coded Inference System for Straggler Mitigation
by: Liu, Xing, et al.
Published: (2025)
by: Liu, Xing, et al.
Published: (2025)
Microarchitectural comparison and in-core modeling of state-of-the-art CPUs: Grace, Sapphire Rapids, and Genoa
by: Laukemann, Jan, et al.
Published: (2024)
by: Laukemann, Jan, et al.
Published: (2024)
Predictive-LoRA: A Proactive and Fragmentation-Aware Serverless Inference System for LLMs
by: Ni, Yinan, et al.
Published: (2025)
by: Ni, Yinan, et al.
Published: (2025)
AnchorTP: Resilient LLM Inference with State-Preserving Elastic Tensor Parallelism
by: Xu, Wendong, et al.
Published: (2025)
by: Xu, Wendong, et al.
Published: (2025)
CloverLeaf on Intel Multi-Core CPUs: A Case Study in Write-Allocate Evasion
by: Laukemann, Jan, et al.
Published: (2023)
by: Laukemann, Jan, et al.
Published: (2023)
THEAS: Efficient Power Management in Multi-Core CPUs via Cache-Aware Resource Scheduling
by: Muhammad, Said, et al.
Published: (2025)
by: Muhammad, Said, et al.
Published: (2025)
Staleness-Centric Optimizations for Parallel Diffusion MoE Inference
by: Luo, Jiajun, et al.
Published: (2024)
by: Luo, Jiajun, et al.
Published: (2024)
Similar Items
-
Inference Performance Optimization for Large Language Models on CPUs
by: He, Pujiang, et al.
Published: (2024) -
Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference
by: Zhang, Haolin, et al.
Published: (2025) -
Inference Acceleration for Large Language Models on CPUs
by: PS, Ditto, et al.
Published: (2024) -
Analyzing the Performance Portability of SYCL across CPUs, GPUs, and Hybrid Systems with SW Sequence Alignment
by: Costanzo, Manuel, et al.
Published: (2024) -
Argus: Token Aware Distributed LLM Inference Optimization
by: Wu, Panlong, et al.
Published: (2025)