ArcLight: A Lightweight LLM Inference Architecture for Many-Core CPUs
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Yuzhuang, Han, Xu, Li, Yuxuan, Che, Wanxiang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Inference Acceleration for Large Language Models on CPUs
by: PS, Ditto, et al.
Published: (2024)
by: PS, Ditto, et al.
Published: (2024)
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers
by: Lu, Yao, et al.
Published: (2026)
by: Lu, Yao, et al.
Published: (2026)
Distributed Inference Performance Optimization for LLMs on CPUs
by: He, Pujiang, et al.
Published: (2024)
by: He, Pujiang, et al.
Published: (2024)
Many Hands Make Light Work: Accelerating Edge Inference via Multi-Client Collaborative Caching
by: Liang, Wenyi, et al.
Published: (2024)
by: Liang, Wenyi, et al.
Published: (2024)
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
by: Liao, Junhan, et al.
Published: (2025)
by: Liao, Junhan, et al.
Published: (2025)
Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference
by: Zhang, Haolin, et al.
Published: (2025)
by: Zhang, Haolin, et al.
Published: (2025)
ScalableHD: Scalable and High-Throughput Hyperdimensional Computing Inference on Multi-Core CPUs
by: Parikh, Dhruv, et al.
Published: (2025)
by: Parikh, Dhruv, et al.
Published: (2025)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
by: Xu, Chuhao, et al.
Published: (2025)
by: Xu, Chuhao, et al.
Published: (2025)
CloverLeaf on Intel Multi-Core CPUs: A Case Study in Write-Allocate Evasion
by: Laukemann, Jan, et al.
Published: (2023)
by: Laukemann, Jan, et al.
Published: (2023)
Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
by: Wu, Yu, et al.
Published: (2025)
by: Wu, Yu, et al.
Published: (2025)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
by: He, Yiyuan, et al.
Published: (2024)
by: He, Yiyuan, et al.
Published: (2024)
Cloud Native System for LLM Inference Serving
by: Xu, Minxian, et al.
Published: (2025)
by: Xu, Minxian, et al.
Published: (2025)
THEAS: Efficient Power Management in Multi-Core CPUs via Cache-Aware Resource Scheduling
by: Muhammad, Said, et al.
Published: (2025)
by: Muhammad, Said, et al.
Published: (2025)
Contemplating a Lightweight Communication Interface for Asynchronous Many-Task Systems
by: Yan, Jiakun, et al.
Published: (2025)
by: Yan, Jiakun, et al.
Published: (2025)
Algebraic Temporal Blocking for Sparse Iterative Solvers on Multi-Core CPUs
by: Alappat, Christie, et al.
Published: (2023)
by: Alappat, Christie, et al.
Published: (2023)
MAGNUS: Generating Data Locality to Accelerate Sparse Matrix-Matrix Multiplication on CPUs
by: Wolfson-Pou, Jordi, et al.
Published: (2025)
by: Wolfson-Pou, Jordi, et al.
Published: (2025)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
by: Hu, Cunchen, et al.
Published: (2024)
by: Hu, Cunchen, et al.
Published: (2024)
A dynamic parallel method for performance optimization on hybrid CPUs
by: Yu, Luo, et al.
Published: (2024)
by: Yu, Luo, et al.
Published: (2024)
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
by: Wei, Jinhui, et al.
Published: (2025)
by: Wei, Jinhui, et al.
Published: (2025)
Analyzing the Performance Portability of SYCL across CPUs, GPUs, and Hybrid Systems with SW Sequence Alignment
by: Costanzo, Manuel, et al.
Published: (2024)
by: Costanzo, Manuel, et al.
Published: (2024)
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
by: Qian, Yulei, et al.
Published: (2024)
by: Qian, Yulei, et al.
Published: (2024)
TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM Inference
by: Zhang, Hongbin, et al.
Published: (2025)
by: Zhang, Hongbin, et al.
Published: (2025)
Token Level Routing Inference System for Edge Devices
by: She, Jianshu, et al.
Published: (2025)
by: She, Jianshu, et al.
Published: (2025)
Aging-aware CPU Core Management for Embodied Carbon Amortization in Cloud LLM Inference
by: Hewage, Tharindu B., et al.
Published: (2025)
by: Hewage, Tharindu B., et al.
Published: (2025)
Efficiently Executing High-throughput Lightweight LLM Inference Applications on Heterogeneous Opportunistic GPU Clusters with Pervasive Context Management
by: Phung, Thanh Son, et al.
Published: (2025)
by: Phung, Thanh Son, et al.
Published: (2025)
InstCache: A Predictive Cache for LLM Serving
by: Zou, Longwei, et al.
Published: (2024)
by: Zou, Longwei, et al.
Published: (2024)
LSRAM: A Lightweight Autoscaling and SLO Resource Allocation Framework for Microservices Based on Gradient Descent
by: Hu, Kan, et al.
Published: (2024)
by: Hu, Kan, et al.
Published: (2024)
Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization
by: Chen, Jiu, et al.
Published: (2026)
by: Chen, Jiu, et al.
Published: (2026)
LeMix: Unified Scheduling for LLM Training and Inference on Multi-GPU Systems
by: Li, Yufei, et al.
Published: (2025)
by: Li, Yufei, et al.
Published: (2025)
Scaling LLM Inference Beyond Amdahl`s Limits via Eliminating Non-Scalable Overheads
by: Zhao, Alan, et al.
Published: (2026)
by: Zhao, Alan, et al.
Published: (2026)
Horizon-LM: A RAM-Centric Architecture for LLM Training
by: Yuan, Zhengqing, et al.
Published: (2026)
by: Yuan, Zhengqing, et al.
Published: (2026)
CIR: Lightweight Container Image for Cross-Platform Deployment
by: Li, Fengzhi, et al.
Published: (2026)
by: Li, Fengzhi, et al.
Published: (2026)
Folding Tensor and Sequence Parallelism for Memory-Efficient Transformer Training & Inference
by: Shyam, Vasu, et al.
Published: (2026)
by: Shyam, Vasu, et al.
Published: (2026)
DeInfer: Efficient Parallel Inferencing for Decomposed Large Language Models
by: Huang, You-Liang, et al.
Published: (2026)
by: Huang, You-Liang, et al.
Published: (2026)
Scaling All-to-all Operations Across Emerging Many-Core Supercomputers
by: Kinkead, Shannon, et al.
Published: (2026)
by: Kinkead, Shannon, et al.
Published: (2026)
LLM-Pilot: Characterize and Optimize Performance of your LLM Inference Services
by: Łazuka, Małgorzata, et al.
Published: (2024)
by: Łazuka, Małgorzata, et al.
Published: (2024)
Efficient Multi-round LLM Inference over Disaggregated Serving
by: He, Wenhao, et al.
Published: (2026)
by: He, Wenhao, et al.
Published: (2026)
PAT: Accelerating LLM Decoding via Prefix-Aware Attention with Resource Efficient Multi-Tile Kernel
by: Yi, Jinjun, et al.
Published: (2025)
by: Yi, Jinjun, et al.
Published: (2025)
Comparison of Vectorization Capabilities of Different Compilers for X86 and ARM CPUs
by: Sakib, Nazmus, et al.
Published: (2025)
by: Sakib, Nazmus, et al.
Published: (2025)
Towards High-Performance and Portable Molecular Docking on CPUs through Vectorization
by: Accordi, Gianmarco, et al.
Published: (2025)
by: Accordi, Gianmarco, et al.
Published: (2025)
Similar Items
-
Inference Acceleration for Large Language Models on CPUs
by: PS, Ditto, et al.
Published: (2024) -
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers
by: Lu, Yao, et al.
Published: (2026) -
Distributed Inference Performance Optimization for LLMs on CPUs
by: He, Pujiang, et al.
Published: (2024) -
Many Hands Make Light Work: Accelerating Edge Inference via Multi-Client Collaborative Caching
by: Liang, Wenyi, et al.
Published: (2024) -
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
by: Liao, Junhan, et al.
Published: (2025)