Fast On-device LLM Inference with NPUs
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Daliang, Zhang, Hao, Yang, Liming, Liu, Ruiqi, Huang, Gang, Xu, Mengwei, Liu, Xuanzhe |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference
by: Yin, Wangsong, et al.
Published: (2025)
by: Yin, Wangsong, et al.
Published: (2025)
Elastic On-Device LLM Service
by: Yin, Wangsong, et al.
Published: (2024)
by: Yin, Wangsong, et al.
Published: (2024)
Quant.npu: Enabling Efficient Mobile NPU Inference for on-device LLMs via Fully Static Quantization
by: Zhang, Jinghe, et al.
Published: (2026)
by: Zhang, Jinghe, et al.
Published: (2026)
MobiEdit: Resource-efficient Knowledge Editing for Personalized On-device LLMs
by: Lu, Zhenyan, et al.
Published: (2025)
by: Lu, Zhenyan, et al.
Published: (2025)
NITRO: LLM Inference on Intel Laptop NPUs
by: Fei, Anthony, et al.
Published: (2024)
by: Fei, Anthony, et al.
Published: (2024)
A-IO: Adaptive Inference Orchestration for Memory-Bound NPUs
by: Zhang, Chen, et al.
Published: (2026)
by: Zhang, Chen, et al.
Published: (2026)
A Survey of Resource-efficient LLM and Multimodal Foundation Models
by: Xu, Mengwei, et al.
Published: (2024)
by: Xu, Mengwei, et al.
Published: (2024)
When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference
by: Li, Pu, et al.
Published: (2026)
by: Li, Pu, et al.
Published: (2026)
BAMAS: Structuring Budget-Aware Multi-Agent Systems
by: Yang, Liming, et al.
Published: (2025)
by: Yang, Liming, et al.
Published: (2025)
Every Software as an Agent: Blueprint and Case Study
by: Xu, Mengwei
Published: (2025)
by: Xu, Mengwei
Published: (2025)
FastQuery: Communication-efficient Embedding Table Query for Private LLM Inference
by: Lin, Chenqi, et al.
Published: (2024)
by: Lin, Chenqi, et al.
Published: (2024)
DroidCall: A Dataset for LLM-powered Android Intent Invocation
by: Xie, Weikai, et al.
Published: (2024)
by: Xie, Weikai, et al.
Published: (2024)
Does Chain-of-Thought Reasoning Help Mobile GUI Agent? An Empirical Study
by: Zhang, Li, et al.
Published: (2025)
by: Zhang, Li, et al.
Published: (2025)
Unleashing Low-Bit Inference on Ascend NPUs: A Comprehensive Evaluation of HiFloat Formats
by: Zhao, Pengxiang, et al.
Published: (2026)
by: Zhao, Pengxiang, et al.
Published: (2026)
Pangu Ultra: Pushing the Limits of Dense Large Language Models on Ascend NPUs
by: Yin, Yichun, et al.
Published: (2025)
by: Yin, Yichun, et al.
Published: (2025)
FwdLLM: Efficient FedLLM using Forward Gradient
by: Xu, Mengwei, et al.
Published: (2023)
by: Xu, Mengwei, et al.
Published: (2023)
Bag of Tricks for Inference-time Computation of LLM Reasoning
by: Liu, Fan, et al.
Published: (2025)
by: Liu, Fan, et al.
Published: (2025)
NVR: Vector Runahead on NPUs for Sparse Memory Access
by: Wang, Hui, et al.
Published: (2025)
by: Wang, Hui, et al.
Published: (2025)
EvoEmo: Towards Evolved Emotional Policies for Adversarial LLM Agents in Multi-Turn Price Negotiation
by: Long, Yunbo, et al.
Published: (2025)
by: Long, Yunbo, et al.
Published: (2025)
LLM as a System Service on Mobile Devices
by: Yin, Wangsong, et al.
Published: (2024)
by: Yin, Wangsong, et al.
Published: (2024)
EmoMAS: Emotion-Aware Multi-Agent System for High-Stakes Edge-Deployable Negotiation with Bayesian Orchestration
by: Long, Yunbo, et al.
Published: (2026)
by: Long, Yunbo, et al.
Published: (2026)
Proceedings Sixth International Workshop on Formal Methods for Autonomous Systems
by: Luckcuck, Matt, et al.
Published: (2024)
by: Luckcuck, Matt, et al.
Published: (2024)
GUI-Shift: Enhancing VLM-Based GUI Agents through Self-supervised Reinforcement Learning
by: Gao, Longxi, et al.
Published: (2025)
by: Gao, Longxi, et al.
Published: (2025)
Fast and Accurate Probing of In-Training LLMs' Downstream Performances
by: Liu, Zhichen, et al.
Published: (2026)
by: Liu, Zhichen, et al.
Published: (2026)
QaRL: Rollout-Aligned Quantization-Aware RL for Fast and Stable Training under Training--Inference Mismatch
by: Gu, Hao, et al.
Published: (2026)
by: Gu, Hao, et al.
Published: (2026)
IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference
by: Yang, Xintong, et al.
Published: (2026)
by: Yang, Xintong, et al.
Published: (2026)
DOCUEVAL: An LLM-based AI Engineering Tool for Building Customisable Document Evaluation Workflows
by: Zhang, Hao, et al.
Published: (2025)
by: Zhang, Hao, et al.
Published: (2025)
Early Diagnosis of Wasted Computation in Multi-Agent LLM Systems via Failure-Aware Observability
by: Li, Xianyou, et al.
Published: (2026)
by: Li, Xianyou, et al.
Published: (2026)
Neuralink: Fast LLM Inference on Smartphones with Neuron Co-Activation Linking
by: Wang, Tuowei, et al.
Published: (2024)
by: Wang, Tuowei, et al.
Published: (2024)
HiFloat4 Format for Language Model Pre-training on Ascend NPUs
by: Taghian, Mehran, et al.
Published: (2026)
by: Taghian, Mehran, et al.
Published: (2026)
MindVL: Towards Efficient and Effective Training of Multimodal Large Language Models on Ascend NPUs
by: Chen, Feilong, et al.
Published: (2025)
by: Chen, Feilong, et al.
Published: (2025)
GADT: Enhancing Transferable Adversarial Attacks through Gradient-guided Adversarial Data Transformation
by: Ma, Yating, et al.
Published: (2024)
by: Ma, Yating, et al.
Published: (2024)
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
by: Liu, Xing, et al.
Published: (2025)
by: Liu, Xing, et al.
Published: (2025)
On the Representation of Pairwise Causal Background Knowledge and Its Applications in Causal Inference
by: Fang, Zhuangyan, et al.
Published: (2022)
by: Fang, Zhuangyan, et al.
Published: (2022)
SemaCDR: LLM-Powered Transferable Semantics for Cross-Domain Sequential Recommendation
by: Zhang, Chunxu, et al.
Published: (2026)
by: Zhang, Chunxu, et al.
Published: (2026)
SynerDiff: Synergetic Continuous Batching for Fast and Parallel Diffusion Model Inference
by: Zhou, Ziqi, et al.
Published: (2026)
by: Zhou, Ziqi, et al.
Published: (2026)
AskNearby: An LLM-Based Application for Neighborhood Information Retrieval and Personalized Cognitive-Map Recommendations
by: Niu, Luyao, et al.
Published: (2025)
by: Niu, Luyao, et al.
Published: (2025)
MISA: Mixture of Indexer Sparse Attention for Long-Context LLM Inference
by: Zhou, Ruijie, et al.
Published: (2026)
by: Zhou, Ruijie, et al.
Published: (2026)
Hexagon-MLIR: An AI Compilation Stack For Qualcomm's Neural Processing Units (NPUs)
by: Absar, Mohammed Javed, et al.
Published: (2026)
by: Absar, Mohammed Javed, et al.
Published: (2026)
FiCoTS: Fine-to-Coarse LLM-Enhanced Hierarchical Cross-Modality Interaction for Time Series Forecasting
by: Lyu, Yafei, et al.
Published: (2025)
by: Lyu, Yafei, et al.
Published: (2025)
Similar Items
-
ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference
by: Yin, Wangsong, et al.
Published: (2025) -
Elastic On-Device LLM Service
by: Yin, Wangsong, et al.
Published: (2024) -
Quant.npu: Enabling Efficient Mobile NPU Inference for on-device LLMs via Fully Static Quantization
by: Zhang, Jinghe, et al.
Published: (2026) -
MobiEdit: Resource-efficient Knowledge Editing for Personalized On-device LLMs
by: Lu, Zhenyan, et al.
Published: (2025) -
NITRO: LLM Inference on Intel Laptop NPUs
by: Fei, Anthony, et al.
Published: (2024)