EasyQuant: An Efficient Data-free Quantization Algorithm for LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tang, Hanlin, Sun, Yifu, Wu, Decheng, Liu, Kai, Zhu, Jianchen, Kang, Zhanhui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Tequila: Trapping-free Ternary Quantization for Large Language Models
von: Huang, Hong, et al.
Veröffentlicht: (2025)
von: Huang, Hong, et al.
Veröffentlicht: (2025)
Sherry: Hardware-Efficient 1.25-Bit Ternary Quantization via Fine-grained Sparsification
von: Huang, Hong, et al.
Veröffentlicht: (2026)
von: Huang, Hong, et al.
Veröffentlicht: (2026)
E-Sparse: Boosting the Large Language Model Inference through Entropy-based N:M Sparsity
von: Li, Yun, et al.
Veröffentlicht: (2023)
von: Li, Yun, et al.
Veröffentlicht: (2023)
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
von: Yan, Xianglong, et al.
Veröffentlicht: (2026)
von: Yan, Xianglong, et al.
Veröffentlicht: (2026)
Quant.npu: Enabling Efficient Mobile NPU Inference for on-device LLMs via Fully Static Quantization
von: Zhang, Jinghe, et al.
Veröffentlicht: (2026)
von: Zhang, Jinghe, et al.
Veröffentlicht: (2026)
SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
von: Zhao, Zhixiong, et al.
Veröffentlicht: (2025)
von: Zhao, Zhixiong, et al.
Veröffentlicht: (2025)
Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge
von: Shen, Xuan, et al.
Veröffentlicht: (2023)
von: Shen, Xuan, et al.
Veröffentlicht: (2023)
NestQuant: Nested Lattice Quantization for Matrix Products and LLMs
von: Savkin, Semyon, et al.
Veröffentlicht: (2025)
von: Savkin, Semyon, et al.
Veröffentlicht: (2025)
SplitQuantV2: Enhancing Low-Bit Quantization of LLMs Without GPUs
von: Song, Jaewoo, et al.
Veröffentlicht: (2025)
von: Song, Jaewoo, et al.
Veröffentlicht: (2025)
DAQ: Delta-Aware Quantization for Post-Training LLM Weight Compression
von: Yu, Xiaoming, et al.
Veröffentlicht: (2026)
von: Yu, Xiaoming, et al.
Veröffentlicht: (2026)
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2022)
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2022)
PolarQuant: Quantizing KV Caches with Polar Transformation
von: Han, Insu, et al.
Veröffentlicht: (2025)
von: Han, Insu, et al.
Veröffentlicht: (2025)
Stem: Rethinking Causal Information Flow in Sparse Attention
von: Niu, Lin, et al.
Veröffentlicht: (2026)
von: Niu, Lin, et al.
Veröffentlicht: (2026)
Easy Samples Are All You Need: Self-Evolving LLMs via Data-Efficient Reinforcement Learning
von: Yu, Zhiyin, et al.
Veröffentlicht: (2026)
von: Yu, Zhiyin, et al.
Veröffentlicht: (2026)
CrossQuant: A Post-Training Quantization Method with Smaller Quantization Kernel for Precise Large Language Model Compression
von: Liu, Wenyuan, et al.
Veröffentlicht: (2024)
von: Liu, Wenyuan, et al.
Veröffentlicht: (2024)
InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization
von: Li, Ke, et al.
Veröffentlicht: (2026)
von: Li, Ke, et al.
Veröffentlicht: (2026)
SplitQuant: Layer Splitting for Low-Bit Neural Network Quantization
von: Song, Jaewoo, et al.
Veröffentlicht: (2025)
von: Song, Jaewoo, et al.
Veröffentlicht: (2025)
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2025)
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2025)
On-the-Fly Adaptation to Quantization: Configuration-Aware LoRA for Efficient Fine-Tuning of Quantized LLMs
von: Ye, Rongguang, et al.
Veröffentlicht: (2025)
von: Ye, Rongguang, et al.
Veröffentlicht: (2025)
QuantLRM: Quantization of Large Reasoning Models via Fine-Tuning Signals
von: Zhang, Nan, et al.
Veröffentlicht: (2026)
von: Zhang, Nan, et al.
Veröffentlicht: (2026)
SmartQuant: CXL-based AI Model Store in Support of Runtime Configurable Weight Quantization
von: Xie, Rui, et al.
Veröffentlicht: (2024)
von: Xie, Rui, et al.
Veröffentlicht: (2024)
Budget-aware Auto Optimizer Configurator
von: Liu, Kang, et al.
Veröffentlicht: (2026)
von: Liu, Kang, et al.
Veröffentlicht: (2026)
Exact Dual Geometry of SOC-ICNN Value Functions
von: Liu, Kang, et al.
Veröffentlicht: (2026)
von: Liu, Kang, et al.
Veröffentlicht: (2026)
MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM
von: Wang, Dongwei, et al.
Veröffentlicht: (2026)
von: Wang, Dongwei, et al.
Veröffentlicht: (2026)
GoQuant: Geometric Orthogonal Residual Projection for Multiplier-Free Power-of-Two Transformer Quantization
von: Xiang, Maoyang, et al.
Veröffentlicht: (2026)
von: Xiang, Maoyang, et al.
Veröffentlicht: (2026)
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
von: Xu, Zukang, et al.
Veröffentlicht: (2025)
von: Xu, Zukang, et al.
Veröffentlicht: (2025)
OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting
von: Hu, Xing, et al.
Veröffentlicht: (2025)
von: Hu, Xing, et al.
Veröffentlicht: (2025)
Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
von: Zhang, Tianao, et al.
Veröffentlicht: (2025)
von: Zhang, Tianao, et al.
Veröffentlicht: (2025)
RaanA: A Fast, Flexible, and Data-Efficient Post-Training Quantization Algorithm
von: Yang, Yongyi, et al.
Veröffentlicht: (2025)
von: Yang, Yongyi, et al.
Veröffentlicht: (2025)
AngelSlim: A more accessible, comprehensive, and efficient toolkit for large model compression
von: Cen, Rui, et al.
Veröffentlicht: (2026)
von: Cen, Rui, et al.
Veröffentlicht: (2026)
FibQuant: Universal Vector Quantization for Random-Access KV-Cache Compression
von: Lee, Namyoon, et al.
Veröffentlicht: (2026)
von: Lee, Namyoon, et al.
Veröffentlicht: (2026)
PrismQuant: Rate-Distortion-Optimal Vector Quantization for Gaussian-Mixture Sources
von: Park, Bumsu, et al.
Veröffentlicht: (2026)
von: Park, Bumsu, et al.
Veröffentlicht: (2026)
QuantMoE-Bench: Examining Post-Training Quantization for Mixture-of-Experts
von: Li, Pingzhi, et al.
Veröffentlicht: (2024)
von: Li, Pingzhi, et al.
Veröffentlicht: (2024)
EEG-DCNet: A Fast and Accurate MI-EEG Dilated CNN Classification Method
von: Peng, Wei, et al.
Veröffentlicht: (2024)
von: Peng, Wei, et al.
Veröffentlicht: (2024)
CleanDiffuser: An Easy-to-use Modularized Library for Diffusion Models in Decision Making
von: Dong, Zibin, et al.
Veröffentlicht: (2024)
von: Dong, Zibin, et al.
Veröffentlicht: (2024)
Evaluating the Generalization Ability of Quantized LLMs: Benchmark, Analysis, and Toolbox
von: Liu, Yijun, et al.
Veröffentlicht: (2024)
von: Liu, Yijun, et al.
Veröffentlicht: (2024)
Easy Problems That LLMs Get Wrong
von: Williams, Sean, et al.
Veröffentlicht: (2024)
von: Williams, Sean, et al.
Veröffentlicht: (2024)
DiscQuant: A Quantization Method for Neural Networks Inspired by Discrepancy Theory
von: Chee, Jerry, et al.
Veröffentlicht: (2025)
von: Chee, Jerry, et al.
Veröffentlicht: (2025)
LogQuant: Log-Distributed 2-Bit Quantization of KV Cache with Superior Accuracy Preservation
von: Chen, Han, et al.
Veröffentlicht: (2025)
von: Chen, Han, et al.
Veröffentlicht: (2025)
ButterflyQuant: Ultra-low-bit LLM Quantization through Learnable Orthogonal Butterfly Transforms
von: Xu, Bingxin, et al.
Veröffentlicht: (2025)
von: Xu, Bingxin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Tequila: Trapping-free Ternary Quantization for Large Language Models
von: Huang, Hong, et al.
Veröffentlicht: (2025) -
Sherry: Hardware-Efficient 1.25-Bit Ternary Quantization via Fine-grained Sparsification
von: Huang, Hong, et al.
Veröffentlicht: (2026) -
E-Sparse: Boosting the Large Language Model Inference through Entropy-based N:M Sparsity
von: Li, Yun, et al.
Veröffentlicht: (2023) -
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
von: Yan, Xianglong, et al.
Veröffentlicht: (2026) -
Quant.npu: Enabling Efficient Mobile NPU Inference for on-device LLMs via Fully Static Quantization
von: Zhang, Jinghe, et al.
Veröffentlicht: (2026)