ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Yin, Wangsong, Xu, Daliang, Xu, Mengwei, Huang, Gang, Liu, Xuanzhe |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Quant.npu: Enabling Efficient Mobile NPU Inference for on-device LLMs via Fully Static Quantization
by: Zhang, Jinghe, et al.
Published: (2026)
by: Zhang, Jinghe, et al.
Published: (2026)
Elastic On-Device LLM Service
by: Yin, Wangsong, et al.
Published: (2024)
by: Yin, Wangsong, et al.
Published: (2024)
Reducing Latency of LLM Search Agent via Speculation-based Algorithm-System Co-Design
by: Huang, Zixiao, et al.
Published: (2025)
by: Huang, Zixiao, et al.
Published: (2025)
Edge Deployment of Small Language Models, a comprehensive comparison of CPU, GPU and NPU backends
by: Prieto, Pablo, et al.
Published: (2025)
by: Prieto, Pablo, et al.
Published: (2025)
Neural Architecture Search of Hybrid Models for NPU-CIM Heterogeneous AR/VR Devices
by: Zhao, Yiwei, et al.
Published: (2024)
by: Zhao, Yiwei, et al.
Published: (2024)
Fast On-device LLM Inference with NPUs
by: Xu, Daliang, et al.
Published: (2024)
by: Xu, Daliang, et al.
Published: (2024)
Neuralink: Fast LLM Inference on Smartphones with Neuron Co-Activation Linking
by: Wang, Tuowei, et al.
Published: (2024)
by: Wang, Tuowei, et al.
Published: (2024)
AscendOptimizer: Episodic Agent for Ascend NPU Operator Optimization
by: Wu, Jiehao, et al.
Published: (2026)
by: Wu, Jiehao, et al.
Published: (2026)
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
by: Lin, Yujun, et al.
Published: (2024)
by: Lin, Yujun, et al.
Published: (2024)
Online Pseudo-average Shifting Attention(PASA) for Robust Low-precision LLM Inference: Algorithms and Numerical Analysis
by: Cheng, Long, et al.
Published: (2025)
by: Cheng, Long, et al.
Published: (2025)
Implementing Keyword Spotting on the MCUX947 Microcontroller with Integrated NPU
by: Jakuš, Petar, et al.
Published: (2025)
by: Jakuš, Petar, et al.
Published: (2025)
AscendCraft: Automatic Ascend NPU Kernel Generation via DSL-Guided Transcompilation
by: Wen, Zhongzhen, et al.
Published: (2026)
by: Wen, Zhongzhen, et al.
Published: (2026)
Pushing the Envelope of LLM Inference on AI-PC and Intel GPUs
by: Georganas, Evangelos, et al.
Published: (2025)
by: Georganas, Evangelos, et al.
Published: (2025)
LLM Swiss Round: Aggregating Multi-Benchmark Performance via Competitive Swiss-System Dynamics
by: Liu, Jiashuo, et al.
Published: (2025)
by: Liu, Jiashuo, et al.
Published: (2025)
Evaluating the Energy Efficiency of NPU-Accelerated Machine Learning Inference on Embedded Microcontrollers
by: Fanariotis, Anastasios, et al.
Published: (2025)
by: Fanariotis, Anastasios, et al.
Published: (2025)
Ragged Paged Attention: A High-Performance and Flexible LLM Inference Kernel for TPU
by: Jiang, Jevin, et al.
Published: (2026)
by: Jiang, Jevin, et al.
Published: (2026)
GreedySnake: Accelerating SSD-Offloaded LLM Training with Efficient Scheduling and Optimizer Step Overlapping
by: Yin, Yishu, et al.
Published: (2025)
by: Yin, Yishu, et al.
Published: (2025)
Private LLM Inference on Consumer Blackwell GPUs: A Practical Guide for Cost-Effective Local Deployment in SMEs
by: Knoop, Jonathan, et al.
Published: (2026)
by: Knoop, Jonathan, et al.
Published: (2026)
eIQ Neutron: Redefining Edge-AI Inference with Integrated NPU and Compiler Innovations
by: Bamberg, Lennart, et al.
Published: (2025)
by: Bamberg, Lennart, et al.
Published: (2025)
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
by: Chu, Kexin, et al.
Published: (2025)
by: Chu, Kexin, et al.
Published: (2025)
Implementation and Optimization of HQC Decoding on NPU-Integrated Devices
by: Chau, Vu Minh, et al.
Published: (2026)
by: Chau, Vu Minh, et al.
Published: (2026)
LLM as a System Service on Mobile Devices
by: Yin, Wangsong, et al.
Published: (2024)
by: Yin, Wangsong, et al.
Published: (2024)
On the Sustainability of AI Inferences in the Edge
by: Sobhani, Ghazal, et al.
Published: (2025)
by: Sobhani, Ghazal, et al.
Published: (2025)
Research on Low-Latency Inference and Training Efficiency Optimization for Graph Neural Network and Large Language Model-Based Recommendation Systems
by: Zhao, Yushang, et al.
Published: (2025)
by: Zhao, Yushang, et al.
Published: (2025)
Towards Efficient Multi-Scale Deformable Attention on NPU
by: Huang, Chenghuan, et al.
Published: (2025)
by: Huang, Chenghuan, et al.
Published: (2025)
A Case Study of Selected PTQ Baselines for Reasoning LLMs on Ascend NPU
by: Luo, Yuchen, et al.
Published: (2026)
by: Luo, Yuchen, et al.
Published: (2026)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
by: Liu, Guangda, et al.
Published: (2024)
by: Liu, Guangda, et al.
Published: (2024)
Forecasting LLM Inference Performance via Hardware-Agnostic Analytical Modeling
by: Patwari, Rajeev, et al.
Published: (2025)
by: Patwari, Rajeev, et al.
Published: (2025)
Bench360: Benchmarking Local LLM Inference from 360 Degrees
by: Stuhlmann, Linus, et al.
Published: (2025)
by: Stuhlmann, Linus, et al.
Published: (2025)
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
by: Shao, Zishan, et al.
Published: (2025)
by: Shao, Zishan, et al.
Published: (2025)
Energy-Efficient Transformer Inference: Optimization Strategies for Time Series Classification
by: Kermani, Arshia, et al.
Published: (2025)
by: Kermani, Arshia, et al.
Published: (2025)
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference
by: Yao, Feiyu, et al.
Published: (2026)
by: Yao, Feiyu, et al.
Published: (2026)
ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching
by: Zhao, Youpeng, et al.
Published: (2024)
by: Zhao, Youpeng, et al.
Published: (2024)
FlashSVD v1.5: Making Low-Rank Transformers Inference Actually Fast
by: Wu, Wenhao, et al.
Published: (2026)
by: Wu, Wenhao, et al.
Published: (2026)
EDGC: Entropy-driven Dynamic Gradient Compression for Efficient LLM Training
by: Yi, Qingao, et al.
Published: (2025)
by: Yi, Qingao, et al.
Published: (2025)
Anatomizing Deep Learning Inference in Web Browsers
by: Wang, Qipeng, et al.
Published: (2024)
by: Wang, Qipeng, et al.
Published: (2024)
MindSpeed RL: Distributed Dataflow for Scalable and Efficient RL Training on Ascend NPU Cluster
by: Feng, Laingjun, et al.
Published: (2025)
by: Feng, Laingjun, et al.
Published: (2025)
SLiM: One-shot Quantization and Sparsity with Low-rank Approximation for LLM Weight Compression
by: Mozaffari, Mohammad, et al.
Published: (2024)
by: Mozaffari, Mohammad, et al.
Published: (2024)
Enabling Performant and Flexible Model-Internal Observability for LLM Inference
by: Yu, Nengneng, et al.
Published: (2026)
by: Yu, Nengneng, et al.
Published: (2026)
A Survey of Resource-efficient LLM and Multimodal Foundation Models
by: Xu, Mengwei, et al.
Published: (2024)
by: Xu, Mengwei, et al.
Published: (2024)
Similar Items
-
Quant.npu: Enabling Efficient Mobile NPU Inference for on-device LLMs via Fully Static Quantization
by: Zhang, Jinghe, et al.
Published: (2026) -
Elastic On-Device LLM Service
by: Yin, Wangsong, et al.
Published: (2024) -
Reducing Latency of LLM Search Agent via Speculation-based Algorithm-System Co-Design
by: Huang, Zixiao, et al.
Published: (2025) -
Edge Deployment of Small Language Models, a comprehensive comparison of CPU, GPU and NPU backends
by: Prieto, Pablo, et al.
Published: (2025) -
Neural Architecture Search of Hybrid Models for NPU-CIM Heterogeneous AR/VR Devices
by: Zhao, Yiwei, et al.
Published: (2024)