Quant-Trim in Practice: Improved Cross-Platform Low-Bit Deployment on Edge NPUs
Fuente:
arXiv
Saved in:
| Main Authors: | Dhahri, Rayen, Urban, Steffen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Shaving Weights with Occam's Razor: Bayesian Sparsification for Neural Networks Using the Marginal Likelihood
by: Dhahri, Rayen, et al.
Published: (2024)
by: Dhahri, Rayen, et al.
Published: (2024)
FrameQuant: Flexible Low-Bit Quantization for Transformers
by: Adepu, Harshavardhan, et al.
Published: (2024)
by: Adepu, Harshavardhan, et al.
Published: (2024)
SplitQuant: Layer Splitting for Low-Bit Neural Network Quantization
by: Song, Jaewoo, et al.
Published: (2025)
by: Song, Jaewoo, et al.
Published: (2025)
InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization
by: Li, Ke, et al.
Published: (2026)
by: Li, Ke, et al.
Published: (2026)
Unleashing Low-Bit Inference on Ascend NPUs: A Comprehensive Evaluation of HiFloat Formats
by: Zhao, Pengxiang, et al.
Published: (2026)
by: Zhao, Pengxiang, et al.
Published: (2026)
SplitQuantV2: Enhancing Low-Bit Quantization of LLMs Without GPUs
by: Song, Jaewoo, et al.
Published: (2025)
by: Song, Jaewoo, et al.
Published: (2025)
SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
by: Zhao, Zhixiong, et al.
Published: (2025)
by: Zhao, Zhixiong, et al.
Published: (2025)
Benchmarking Ultra-Low-Power $μ$NPUs
by: Millar, Josh, et al.
Published: (2025)
by: Millar, Josh, et al.
Published: (2025)
Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
by: Zhang, Tianao, et al.
Published: (2025)
by: Zhang, Tianao, et al.
Published: (2025)
TruncQuant: Truncation-Ready Quantization for DNNs with Flexible Weight Bit Precision
by: Kim, Jinhee, et al.
Published: (2025)
by: Kim, Jinhee, et al.
Published: (2025)
NanoQuant: Efficient Sub-1-Bit Quantization of Large Language Models
by: Chong, Hyochan, et al.
Published: (2026)
by: Chong, Hyochan, et al.
Published: (2026)
pQuant: Towards Effective Low-Bit Language Models via Decoupled Linear Quantization-Aware Training
by: Zhang, Wenzheng, et al.
Published: (2026)
by: Zhang, Wenzheng, et al.
Published: (2026)
FastAttention: Extend FlashAttention2 to NPUs and Low-resource GPUs
by: Lin, Haoran, et al.
Published: (2024)
by: Lin, Haoran, et al.
Published: (2024)
BitRL: Reinforcement Learning with 1-bit Quantized Language Models for Resource-Constrained Edge Deployment
by: Sajid, Md. Ashiq Ul Islam, et al.
Published: (2026)
by: Sajid, Md. Ashiq Ul Islam, et al.
Published: (2026)
SHIELD: A Segmented Hierarchical Memory Architecture for Energy-Efficient LLM Inference on Edge NPUs
by: Zhang, Jintao, et al.
Published: (2026)
by: Zhang, Jintao, et al.
Published: (2026)
From Principles to Practice: A Systematic Study of LLM Serving on Multi-core NPUs
by: Zhu, Tianhao, et al.
Published: (2025)
by: Zhu, Tianhao, et al.
Published: (2025)
Squat: Quant Small Language Models on the Edge
by: Shen, Xuan, et al.
Published: (2024)
by: Shen, Xuan, et al.
Published: (2024)
1-Bit Wonder: Improving QAT Performance in the Low-Bit Regime through K-Means Quantization
by: Maskey, Sohir, et al.
Published: (2026)
by: Maskey, Sohir, et al.
Published: (2026)
Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization
by: Xi, Haocheng, et al.
Published: (2026)
by: Xi, Haocheng, et al.
Published: (2026)
MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM
by: Wang, Dongwei, et al.
Published: (2026)
by: Wang, Dongwei, et al.
Published: (2026)
Revisiting RaBitQ and TurboQuant: A Symmetric Comparison of Methods, Theory, and Experiments
by: Gao, Jianyang, et al.
Published: (2026)
by: Gao, Jianyang, et al.
Published: (2026)
Efficient Mixture-of-Experts LLM Inference with Apple Silicon NPUs
by: Benazir, Afsara, et al.
Published: (2026)
by: Benazir, Afsara, et al.
Published: (2026)
Towards Real-Time ECG and EMG Modeling on $μ$NPUs
by: Millar, Josh, et al.
Published: (2026)
by: Millar, Josh, et al.
Published: (2026)
Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM Deployment
by: Lee, Deokjae, et al.
Published: (2025)
by: Lee, Deokjae, et al.
Published: (2025)
LogQuant: Log-Distributed 2-Bit Quantization of KV Cache with Superior Accuracy Preservation
by: Chen, Han, et al.
Published: (2025)
by: Chen, Han, et al.
Published: (2025)
Bit-by-Bit: Progressive QAT Strategy with Outlier Channel Splitting for Stable Low-Bit LLMs
by: Xu, Binxing, et al.
Published: (2026)
by: Xu, Binxing, et al.
Published: (2026)
SigmaQuant: Hardware-Aware Heterogeneous Quantization Method for Edge DNN Inference
by: Liu, Qunyou, et al.
Published: (2026)
by: Liu, Qunyou, et al.
Published: (2026)
Edge-ASR: Towards Low-Bit Quantization of Automatic Speech Recognition Models
by: Feng, Chen, et al.
Published: (2025)
by: Feng, Chen, et al.
Published: (2025)
Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge
by: Shen, Xuan, et al.
Published: (2023)
by: Shen, Xuan, et al.
Published: (2023)
Test-Time Adaptation by Causal Trimming
by: Liu, Yingnan, et al.
Published: (2025)
by: Liu, Yingnan, et al.
Published: (2025)
From Bits to Chips: An LLM-based Hardware-Aware Quantization Agent for Streamlined Deployment of LLMs
by: Deng, Kaiyuan, et al.
Published: (2026)
by: Deng, Kaiyuan, et al.
Published: (2026)
EAGLE-Pangu: Accelerator-Safe Tree Speculative Decoding on Ascend NPUs
by: Han, Chang, et al.
Published: (2026)
by: Han, Chang, et al.
Published: (2026)
QuantFL: Sustainable Federated Learning for Edge IoT via Pre-Trained Model Quantisation
by: Herath, Charuka, et al.
Published: (2026)
by: Herath, Charuka, et al.
Published: (2026)
CodeQuant: Unified Clustering and Quantization for Enhanced Outlier Smoothing in Low-Precision Mixture-of-Experts
by: Yin, Xiangyang, et al.
Published: (2026)
by: Yin, Xiangyang, et al.
Published: (2026)
Energy-Efficient Vision Transformer Inference for Edge-AI Deployment
by: Amanzhol, Nursultan, et al.
Published: (2025)
by: Amanzhol, Nursultan, et al.
Published: (2025)
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
by: Yan, Xianglong, et al.
Published: (2026)
by: Yan, Xianglong, et al.
Published: (2026)
Real Image Denoising with Knowledge Distillation for High-Performance Mobile NPUs
by: Kayani, Faraz, et al.
Published: (2026)
by: Kayani, Faraz, et al.
Published: (2026)
LittleBit: Ultra Low-Bit Quantization via Latent Factorization
by: Lee, Banseok, et al.
Published: (2025)
by: Lee, Banseok, et al.
Published: (2025)
Efficient LLaMA-3.2-Vision by Trimming Cross-attended Visual Features
by: Lee, Jewon, et al.
Published: (2025)
by: Lee, Jewon, et al.
Published: (2025)
Training with Fewer Bits: Unlocking Edge LLMs Training with Stochastic Rounding
by: Liu, Taowen, et al.
Published: (2025)
by: Liu, Taowen, et al.
Published: (2025)
Similar Items
-
Shaving Weights with Occam's Razor: Bayesian Sparsification for Neural Networks Using the Marginal Likelihood
by: Dhahri, Rayen, et al.
Published: (2024) -
FrameQuant: Flexible Low-Bit Quantization for Transformers
by: Adepu, Harshavardhan, et al.
Published: (2024) -
SplitQuant: Layer Splitting for Low-Bit Neural Network Quantization
by: Song, Jaewoo, et al.
Published: (2025) -
InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization
by: Li, Ke, et al.
Published: (2026) -
Unleashing Low-Bit Inference on Ascend NPUs: A Comprehensive Evaluation of HiFloat Formats
by: Zhao, Pengxiang, et al.
Published: (2026)