From Principles to Practice: A Systematic Study of LLM Serving on Multi-core NPUs
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Tianhao, Feng, Dahu, Feng, Erhu, Xia, Yubin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SHIELD: A Segmented Hierarchical Memory Architecture for Energy-Efficient LLM Inference on Edge NPUs
by: Zhang, Jintao, et al.
Published: (2026)
by: Zhang, Jintao, et al.
Published: (2026)
Topology-Aware Virtualization over Inter-Core Connected Neural Processing Units
by: Feng, Dahu, et al.
Published: (2025)
by: Feng, Dahu, et al.
Published: (2025)
Benchmarking Ultra-Low-Power $μ$NPUs
by: Millar, Josh, et al.
Published: (2025)
by: Millar, Josh, et al.
Published: (2025)
Pimba: A Processing-in-Memory Acceleration for Post-Transformer Large Language Model Serving
by: Kim, Wonung, et al.
Published: (2025)
by: Kim, Wonung, et al.
Published: (2025)
From Loop Nests to Silicon: Mapping AI Workloads onto AMD NPUs with MLIR-AIR
by: Wang, Erwei, et al.
Published: (2025)
by: Wang, Erwei, et al.
Published: (2025)
L3: DIMM-PIM Integrated Architecture and Coordination for Scalable Long-Context LLM Inference
by: Liu, Qingyuan, et al.
Published: (2025)
by: Liu, Qingyuan, et al.
Published: (2025)
Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization
by: Kim, Minsu, et al.
Published: (2025)
by: Kim, Minsu, et al.
Published: (2025)
ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs
by: Yang, Jinwu, et al.
Published: (2026)
by: Yang, Jinwu, et al.
Published: (2026)
COMET: Towards Partical W4A4KV4 LLMs Serving
by: Liu, Lian, et al.
Published: (2024)
by: Liu, Lian, et al.
Published: (2024)
PASCAL: A Phase-Aware Scheduling Algorithm for Serving Reasoning-based Large Language Models
by: Cho, Eunyeong, et al.
Published: (2026)
by: Cho, Eunyeong, et al.
Published: (2026)
MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving
by: Lee, Jungi, et al.
Published: (2025)
by: Lee, Jungi, et al.
Published: (2025)
VeriCache: Turning Lossy KV Cache into Lossless LLM Inference
by: Yao, Jiayi, et al.
Published: (2026)
by: Yao, Jiayi, et al.
Published: (2026)
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference
by: Yubeaton, Patrick, et al.
Published: (2025)
by: Yubeaton, Patrick, et al.
Published: (2025)
Bit-Flip Vulnerability of Shared KV-Cache Blocks in LLM Serving Systems
by: Yamamoto, Yuji, et al.
Published: (2026)
by: Yamamoto, Yuji, et al.
Published: (2026)
FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design
by: Xia, Haojun, et al.
Published: (2024)
by: Xia, Haojun, et al.
Published: (2024)
From LLM to Silicon: RL-Driven ASIC Architecture Exploration for On-Device AI Inference
by: Ganti, Ravindra, et al.
Published: (2026)
by: Ganti, Ravindra, et al.
Published: (2026)
SetupKit: Efficient Multi-Corner Setup/Hold Time Characterization Using Bias-Enhanced Interpolation and Active Learning
by: Zhou, Junzhuo, et al.
Published: (2025)
by: Zhou, Junzhuo, et al.
Published: (2025)
RTLSeek: Boosting the LLM-Based RTL Generation with Multi-Stage Diversity-Oriented Reinforcement Learning
by: Zhang, Xinyu, et al.
Published: (2026)
by: Zhang, Xinyu, et al.
Published: (2026)
EPIM: Efficient Processing-In-Memory Accelerators based on Epitome
by: Wang, Chenyu, et al.
Published: (2023)
by: Wang, Chenyu, et al.
Published: (2023)
FPGA-Enabled Machine Learning Applications in Earth Observation: A Systematic Review
by: Léonard, Cédric, et al.
Published: (2025)
by: Léonard, Cédric, et al.
Published: (2025)
MPM-LLM4DSE: Reaching the Pareto Frontier in HLS with Multimodal Learning and LLM-Driven Exploration
by: Xu, Lei, et al.
Published: (2026)
by: Xu, Lei, et al.
Published: (2026)
CIMFlow: An Integrated Framework for Systematic Design and Evaluation of Digital CIM Architectures
by: Qi, Yingjie, et al.
Published: (2025)
by: Qi, Yingjie, et al.
Published: (2025)
T-MAN: Enabling End-to-End Low-Bit LLM Inference on NPUs via Unified Table Lookup
by: Wei, Jianyu, et al.
Published: (2025)
by: Wei, Jianyu, et al.
Published: (2025)
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats
by: Chen, Yuzong, et al.
Published: (2025)
by: Chen, Yuzong, et al.
Published: (2025)
Designing Efficient LLM Accelerators for Edge Devices
by: Haris, Jude, et al.
Published: (2024)
by: Haris, Jude, et al.
Published: (2024)
KLLM: Fast LLM Inference with K-Means Quantization
by: Wu, Xueying, et al.
Published: (2025)
by: Wu, Xueying, et al.
Published: (2025)
CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUs
by: Cai, Tianhao, et al.
Published: (2025)
by: Cai, Tianhao, et al.
Published: (2025)
Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE Serving
by: Pan, Yue, et al.
Published: (2025)
by: Pan, Yue, et al.
Published: (2025)
TimingLLM: A Two-Stage Retrieval-Augmented Framework for Pre-Synthesis Timing Prediction from Verilog
by: Abdollahi, Armin, et al.
Published: (2026)
by: Abdollahi, Armin, et al.
Published: (2026)
LLM-based AI Agent for Sizing of Analog and Mixed Signal Circuit
by: Liu, Chang, et al.
Published: (2025)
by: Liu, Chang, et al.
Published: (2025)
LLM-USO: Large Language Model-based Universal Sizing Optimizer
by: S, Karthik Somayaji N., et al.
Published: (2025)
by: S, Karthik Somayaji N., et al.
Published: (2025)
FedChip: Federated LLM for Artificial Intelligence Accelerator Chip Design
by: Nazzal, Mahmoud, et al.
Published: (2025)
by: Nazzal, Mahmoud, et al.
Published: (2025)
BitMoD: Bit-serial Mixture-of-Datatype LLM Acceleration
by: Chen, Yuzong, et al.
Published: (2024)
by: Chen, Yuzong, et al.
Published: (2024)
EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture
by: Duan, Bowen, et al.
Published: (2026)
by: Duan, Bowen, et al.
Published: (2026)
Unveiling Environmental Impacts of Large Language Model Serving: A Functional Unit View
by: Wu, Yanran, et al.
Published: (2025)
by: Wu, Yanran, et al.
Published: (2025)
EEsizer: LLM-Based AI Agent for Sizing of Analog and Mixed Signal Circuit
by: Liu, Chang, et al.
Published: (2025)
by: Liu, Chang, et al.
Published: (2025)
When Forgetting Builds Reliability: LLM Unlearning for Reliable Hardware Code Generation
by: Liang, Yiwen, et al.
Published: (2025)
by: Liang, Yiwen, et al.
Published: (2025)
MAGE: A Multi-Agent Engine for Automated RTL Code Generation
by: Zhao, Yujie, et al.
Published: (2024)
by: Zhao, Yujie, et al.
Published: (2024)
T-SAR: A Full-Stack Co-design for CPU-Only Ternary LLM Inference via In-Place SIMD ALU Reorganization
by: Oh, Hyunwoo, et al.
Published: (2025)
by: Oh, Hyunwoo, et al.
Published: (2025)
The prediction of the quality of results in Logic Synthesis using Transformer and Graph Neural Networks
by: Yang, Chenghao, et al.
Published: (2022)
by: Yang, Chenghao, et al.
Published: (2022)
Similar Items
-
SHIELD: A Segmented Hierarchical Memory Architecture for Energy-Efficient LLM Inference on Edge NPUs
by: Zhang, Jintao, et al.
Published: (2026) -
Topology-Aware Virtualization over Inter-Core Connected Neural Processing Units
by: Feng, Dahu, et al.
Published: (2025) -
Benchmarking Ultra-Low-Power $μ$NPUs
by: Millar, Josh, et al.
Published: (2025) -
Pimba: A Processing-in-Memory Acceleration for Post-Transformer Large Language Model Serving
by: Kim, Wonung, et al.
Published: (2025) -
From Loop Nests to Silicon: Mapping AI Workloads onto AMD NPUs with MLIR-AIR
by: Wang, Erwei, et al.
Published: (2025)