Quant.npu: Enabling Efficient Mobile NPU Inference for on-device LLMs via Fully Static Quantization
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Jinghe, Xu, Daliang, Wang, Chenghua, Xie, Weikai, Qi, Tao, Ma, Yun, Xu, Mengwei, Huang, Gang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference
di: Yin, Wangsong, et al.
Pubblicazione: (2025)
di: Yin, Wangsong, et al.
Pubblicazione: (2025)
Accelerating Mobile Language Model via Speculative Decoding and NPU-Coordinated Execution
di: Chen, Zhiyang, et al.
Pubblicazione: (2025)
di: Chen, Zhiyang, et al.
Pubblicazione: (2025)
Fast On-device LLM Inference with NPUs
di: Xu, Daliang, et al.
Pubblicazione: (2024)
di: Xu, Daliang, et al.
Pubblicazione: (2024)
NanoSpec: Accelerating Speculative Decoding using Minimalist In-Context Vocabularies
di: Chen, Zhiyang, et al.
Pubblicazione: (2026)
di: Chen, Zhiyang, et al.
Pubblicazione: (2026)
MobileQuant: Mobile-friendly Quantization for On-device Language Models
di: Tan, Fuwen, et al.
Pubblicazione: (2024)
di: Tan, Fuwen, et al.
Pubblicazione: (2024)
MobiEdit: Resource-efficient Knowledge Editing for Personalized On-device LLMs
di: Lu, Zhenyan, et al.
Pubblicazione: (2025)
di: Lu, Zhenyan, et al.
Pubblicazione: (2025)
Elastic On-Device LLM Service
di: Yin, Wangsong, et al.
Pubblicazione: (2024)
di: Yin, Wangsong, et al.
Pubblicazione: (2024)
PrivQuant: Communication-Efficient Private Inference with Quantized Network/Protocol Co-Optimization
di: Xu, Tianshi, et al.
Pubblicazione: (2024)
di: Xu, Tianshi, et al.
Pubblicazione: (2024)
PhoneLM:an Efficient and Capable Small Language Model Family through Principled Pre-training
di: Yi, Rongjie, et al.
Pubblicazione: (2024)
di: Yi, Rongjie, et al.
Pubblicazione: (2024)
Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge
di: Shen, Xuan, et al.
Pubblicazione: (2023)
di: Shen, Xuan, et al.
Pubblicazione: (2023)
One QuantLLM for ALL: Fine-tuning Quantized LLMs Once for Efficient Deployments
di: Yi, Ke, et al.
Pubblicazione: (2024)
di: Yi, Ke, et al.
Pubblicazione: (2024)
EasyQuant: An Efficient Data-free Quantization Algorithm for LLMs
di: Tang, Hanlin, et al.
Pubblicazione: (2024)
di: Tang, Hanlin, et al.
Pubblicazione: (2024)
ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
di: Liang, Yesheng, et al.
Pubblicazione: (2025)
di: Liang, Yesheng, et al.
Pubblicazione: (2025)
Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs
di: Lu, Haiquan, et al.
Pubblicazione: (2026)
di: Lu, Haiquan, et al.
Pubblicazione: (2026)
QuantV2X: A Fully Quantized Multi-Agent System for Cooperative Perception
di: Zhao, Seth Z., et al.
Pubblicazione: (2025)
di: Zhao, Seth Z., et al.
Pubblicazione: (2025)
Towards Efficient Multi-Scale Deformable Attention on NPU
di: Huang, Chenghuan, et al.
Pubblicazione: (2025)
di: Huang, Chenghuan, et al.
Pubblicazione: (2025)
EfficientQuant: An Efficient Post-Training Quantization for CNN-Transformer Hybrid Models on Edge Devices
di: Saha, Shaibal, et al.
Pubblicazione: (2025)
di: Saha, Shaibal, et al.
Pubblicazione: (2025)
MergeQuant: Accurate 4-bit Static Quantization of Large Language Models by Channel-wise Calibration
di: Wang, Jinguang, et al.
Pubblicazione: (2025)
di: Wang, Jinguang, et al.
Pubblicazione: (2025)
DroidCall: A Dataset for LLM-powered Android Intent Invocation
di: Xie, Weikai, et al.
Pubblicazione: (2024)
di: Xie, Weikai, et al.
Pubblicazione: (2024)
UniCAIM: A Unified CAM/CIM Architecture with Static-Dynamic KV Cache Pruning for Efficient Long-Context LLM Inference
di: Xu, Weikai, et al.
Pubblicazione: (2025)
di: Xu, Weikai, et al.
Pubblicazione: (2025)
CalibQuant: 1-Bit KV Cache Quantization for Multimodal LLMs
di: Han, Insu, et al.
Pubblicazione: (2025)
di: Han, Insu, et al.
Pubblicazione: (2025)
SliderQuant: Accurate Post-Training Quantization for LLMs
di: Wang, Shigeng, et al.
Pubblicazione: (2026)
di: Wang, Shigeng, et al.
Pubblicazione: (2026)
QuantFace: Efficient Quantization for Face Restoration
di: Li, Jiatong, et al.
Pubblicazione: (2025)
di: Li, Jiatong, et al.
Pubblicazione: (2025)
LLMEasyQuant: Scalable Quantization for Parallel and Distributed LLM Inference
di: Liu, Dong, et al.
Pubblicazione: (2024)
di: Liu, Dong, et al.
Pubblicazione: (2024)
DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs
di: Lin, Haokun, et al.
Pubblicazione: (2024)
di: Lin, Haokun, et al.
Pubblicazione: (2024)
NestQuant: Nested Lattice Quantization for Matrix Products and LLMs
di: Savkin, Semyon, et al.
Pubblicazione: (2025)
di: Savkin, Semyon, et al.
Pubblicazione: (2025)
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
di: Yan, Xianglong, et al.
Pubblicazione: (2026)
di: Yan, Xianglong, et al.
Pubblicazione: (2026)
NPU Design for Diffusion Language Model Inference
di: Lou, Binglei, et al.
Pubblicazione: (2026)
di: Lou, Binglei, et al.
Pubblicazione: (2026)
DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization
di: Shao, Yuantian, et al.
Pubblicazione: (2025)
di: Shao, Yuantian, et al.
Pubblicazione: (2025)
GS-Quant: Granular Semantic and Generative Structural Quantization for Knowledge Graph Completion
di: Xie, Qizhuo, et al.
Pubblicazione: (2026)
di: Xie, Qizhuo, et al.
Pubblicazione: (2026)
SigmaQuant: Hardware-Aware Heterogeneous Quantization Method for Edge DNN Inference
di: Liu, Qunyou, et al.
Pubblicazione: (2026)
di: Liu, Qunyou, et al.
Pubblicazione: (2026)
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
di: Xu, Zukang, et al.
Pubblicazione: (2025)
di: Xu, Zukang, et al.
Pubblicazione: (2025)
I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models
di: Hu, Xing, et al.
Pubblicazione: (2024)
di: Hu, Xing, et al.
Pubblicazione: (2024)
LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load
di: Tummalapalli, Pranay, et al.
Pubblicazione: (2026)
di: Tummalapalli, Pranay, et al.
Pubblicazione: (2026)
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization
di: Tao, Wei, et al.
Pubblicazione: (2026)
di: Tao, Wei, et al.
Pubblicazione: (2026)
DilateQuant: Accurate and Efficient Diffusion Quantization via Weight Dilation
di: Liu, Xuewen, et al.
Pubblicazione: (2024)
di: Liu, Xuewen, et al.
Pubblicazione: (2024)
OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models
di: Shao, Wenqi, et al.
Pubblicazione: (2023)
di: Shao, Wenqi, et al.
Pubblicazione: (2023)
Scaling LLM Test-Time Compute with Mobile NPU on Smartphones
di: Hao, Zixu, et al.
Pubblicazione: (2025)
di: Hao, Zixu, et al.
Pubblicazione: (2025)
Does Chain-of-Thought Reasoning Help Mobile GUI Agent? An Empirical Study
di: Zhang, Li, et al.
Pubblicazione: (2025)
di: Zhang, Li, et al.
Pubblicazione: (2025)
Every Software as an Agent: Blueprint and Case Study
di: Xu, Mengwei
Pubblicazione: (2025)
di: Xu, Mengwei
Pubblicazione: (2025)
Documenti analoghi
-
ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference
di: Yin, Wangsong, et al.
Pubblicazione: (2025) -
Accelerating Mobile Language Model via Speculative Decoding and NPU-Coordinated Execution
di: Chen, Zhiyang, et al.
Pubblicazione: (2025) -
Fast On-device LLM Inference with NPUs
di: Xu, Daliang, et al.
Pubblicazione: (2024) -
NanoSpec: Accelerating Speculative Decoding using Minimalist In-Context Vocabularies
di: Chen, Zhiyang, et al.
Pubblicazione: (2026) -
MobileQuant: Mobile-friendly Quantization for On-device Language Models
di: Tan, Fuwen, et al.
Pubblicazione: (2024)