DecDEC: A Systems Approach to Advancing Low-Bit LLM Quantization
Fuente:
arXiv
Saved in:
| Main Authors: | Park, Yeonhong, Hyun, Jake, Kim, Hojoon, Lee, Jae W. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs
by: Park, Yeonhong, et al.
Published: (2024)
by: Park, Yeonhong, et al.
Published: (2024)
GuidedQuant: Large Language Model Quantization via Exploiting End Loss Guidance
by: Kim, Jinuk, et al.
Published: (2025)
by: Kim, Jinuk, et al.
Published: (2025)
MAGE: All-[MASK] Block Already Knows Where to Look in Diffusion LLM
by: Kwon, Omin, et al.
Published: (2026)
by: Kwon, Omin, et al.
Published: (2026)
DP-LLM: Runtime Model Adaptation with Dynamic Layer-wise Precision Assignment
by: Kwon, Sangwoo, et al.
Published: (2025)
by: Kwon, Sangwoo, et al.
Published: (2025)
Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM Deployment
by: Lee, Deokjae, et al.
Published: (2025)
by: Lee, Deokjae, et al.
Published: (2025)
LittleBit: Ultra Low-Bit Quantization via Latent Factorization
by: Lee, Banseok, et al.
Published: (2025)
by: Lee, Banseok, et al.
Published: (2025)
SBVR: Summation of BitVector Representation for Efficient LLM Quantization
by: Bang, Wonjun, et al.
Published: (2025)
by: Bang, Wonjun, et al.
Published: (2025)
Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature Estimate
by: Kim, Jaemin, et al.
Published: (2026)
by: Kim, Jaemin, et al.
Published: (2026)
Learning Grouped Lattice Vector Quantizers for Low-Bit LLM Compression
by: Zhang, Xi, et al.
Published: (2025)
by: Zhang, Xi, et al.
Published: (2025)
InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization
by: Li, Ke, et al.
Published: (2026)
by: Li, Ke, et al.
Published: (2026)
Why Do Some Inputs Break Low-Bit LLM Quantization?
by: Chang, Ting-Yun, et al.
Published: (2025)
by: Chang, Ting-Yun, et al.
Published: (2025)
SERQ: Saliency-Aware Low-Rank Error Reconstruction for LLM Quantization
by: Park, Yeonsik, et al.
Published: (2026)
by: Park, Yeonsik, et al.
Published: (2026)
TesseraQ: Ultra Low-Bit LLM Post-Training Quantization with Block Reconstruction
by: Li, Yuhang, et al.
Published: (2024)
by: Li, Yuhang, et al.
Published: (2024)
MimiQ: Low-Bit Data-Free Quantization of Vision Transformers with Encouraging Inter-Head Attention Similarity
by: Choi, Kanghyun, et al.
Published: (2024)
by: Choi, Kanghyun, et al.
Published: (2024)
BitSnap: Checkpoint Sparsification and Quantization in LLM Training
by: Peng, Yanxin, et al.
Published: (2025)
by: Peng, Yanxin, et al.
Published: (2025)
Technical Report: Activation Residual Hessian Quantization (ARHQ) for Low-Bit LLM Quantization
by: Wang, YiFeng, et al.
Published: (2026)
by: Wang, YiFeng, et al.
Published: (2026)
1-Bit Wonder: Improving QAT Performance in the Low-Bit Regime through K-Means Quantization
by: Maskey, Sohir, et al.
Published: (2026)
by: Maskey, Sohir, et al.
Published: (2026)
FrameQuant: Flexible Low-Bit Quantization for Transformers
by: Adepu, Harshavardhan, et al.
Published: (2024)
by: Adepu, Harshavardhan, et al.
Published: (2024)
TruncQuant: Truncation-Ready Quantization for DNNs with Flexible Weight Bit Precision
by: Kim, Jinhee, et al.
Published: (2025)
by: Kim, Jinhee, et al.
Published: (2025)
CLAQ: Pushing the Limits of Low-Bit Post-Training Quantization for LLMs
by: Wang, Haoyu, et al.
Published: (2024)
by: Wang, Haoyu, et al.
Published: (2024)
Log-Time K-Means Clustering for 1D Data: Novel Approaches with Proof and Implementation
by: Hyun, Jake
Published: (2024)
by: Hyun, Jake
Published: (2024)
MSQ: Memory-Efficient Bit Sparsification Quantization
by: Han, Seokho, et al.
Published: (2025)
by: Han, Seokho, et al.
Published: (2025)
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache
by: Son, Donghyun, et al.
Published: (2025)
by: Son, Donghyun, et al.
Published: (2025)
Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
by: Zhang, Tianao, et al.
Published: (2025)
by: Zhang, Tianao, et al.
Published: (2025)
LoRAQuant: Mixed-Precision Quantization of LoRA to Ultra-Low Bits
by: Mirzaei, Amir Reza, et al.
Published: (2025)
by: Mirzaei, Amir Reza, et al.
Published: (2025)
Can Large Language Models Develop Strategic Reasoning? Post-training Insights from Learning Chess
by: Hwang, Dongyoon, et al.
Published: (2025)
by: Hwang, Dongyoon, et al.
Published: (2025)
NanoQuant: Efficient Sub-1-Bit Quantization of Large Language Models
by: Chong, Hyochan, et al.
Published: (2026)
by: Chong, Hyochan, et al.
Published: (2026)
ABQ-LLM: Arbitrary-Bit Quantized Inference Acceleration for Large Language Models
by: Zeng, Chao, et al.
Published: (2024)
by: Zeng, Chao, et al.
Published: (2024)
BitsMoE: Efficient Spectral Energy-Guided Bit Allocation for MoE LLM Quantization
by: Zhao, Jiayu, et al.
Published: (2026)
by: Zhao, Jiayu, et al.
Published: (2026)
SAW-INT4: System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving
by: Jia, Jinda, et al.
Published: (2026)
by: Jia, Jinda, et al.
Published: (2026)
SplitQuant: Layer Splitting for Low-Bit Neural Network Quantization
by: Song, Jaewoo, et al.
Published: (2025)
by: Song, Jaewoo, et al.
Published: (2025)
FIRE: Frobenius-Isometry Reinitialization for Balancing the Stability-Plasticity Tradeoff
by: Han, Isaac, et al.
Published: (2026)
by: Han, Isaac, et al.
Published: (2026)
Influence-Inspired Spectral Rotations for Extreme Low-Bit LLM Quantization
by: Pavlov, Gorgi
Published: (2026)
by: Pavlov, Gorgi
Published: (2026)
MuonQ: Enhancing Low-Bit Muon Quantization via Directional Fidelity Optimization
by: Su, Yupeng, et al.
Published: (2026)
by: Su, Yupeng, et al.
Published: (2026)
Efficient Multi-bit Quantization Network Training via Weight Bias Correction and Bit-wise Coreset Sampling
by: Kim, Jinhee, et al.
Published: (2025)
by: Kim, Jinhee, et al.
Published: (2025)
I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models
by: Hu, Xing, et al.
Published: (2024)
by: Hu, Xing, et al.
Published: (2024)
Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens
by: Ouyang, Xu, et al.
Published: (2024)
by: Ouyang, Xu, et al.
Published: (2024)
Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models
by: Park, Jungwoo, et al.
Published: (2025)
by: Park, Jungwoo, et al.
Published: (2025)
One-Bit Quantization for Random Features Models
by: Akhtiamov, Danil, et al.
Published: (2025)
by: Akhtiamov, Danil, et al.
Published: (2025)
ECQ$^{\text{x}}$: Explainability-Driven Quantization for Low-Bit and Sparse DNNs
by: Becking, Daniel, et al.
Published: (2021)
by: Becking, Daniel, et al.
Published: (2021)
Similar Items
-
Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs
by: Park, Yeonhong, et al.
Published: (2024) -
GuidedQuant: Large Language Model Quantization via Exploiting End Loss Guidance
by: Kim, Jinuk, et al.
Published: (2025) -
MAGE: All-[MASK] Block Already Knows Where to Look in Diffusion LLM
by: Kwon, Omin, et al.
Published: (2026) -
DP-LLM: Runtime Model Adaptation with Dynamic Layer-wise Precision Assignment
by: Kwon, Sangwoo, et al.
Published: (2025) -
Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM Deployment
by: Lee, Deokjae, et al.
Published: (2025)