Technical Report: Activation Residual Hessian Quantization (ARHQ) for Low-Bit LLM Quantization
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, YiFeng, Sun, Zhun, Sakaguchi, Keisuke |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhancing Post-Training Quantization via Future Activation Awareness
by: Lv, Zheqi, et al.
Published: (2026)
by: Lv, Zheqi, et al.
Published: (2026)
MARR: Module-Adaptive Residual Reconstruction for Low-Bit Post-Training Quantization
by: Su, Le, et al.
Published: (2026)
by: Su, Le, et al.
Published: (2026)
Low-Bit, High-Fidelity: Optimal Transport Quantization for Flow Matching
by: Varam, Dara, et al.
Published: (2025)
by: Varam, Dara, et al.
Published: (2025)
ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization
by: Liu, Zechun, et al.
Published: (2025)
by: Liu, Zechun, et al.
Published: (2025)
BayesQ: Uncertainty-Guided Bayesian Quantization
by: Lamaakal, Ismail, et al.
Published: (2025)
by: Lamaakal, Ismail, et al.
Published: (2025)
QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
MGVQ: Synergizing Multi-dimensional Sensitivity-Aware and Gradient-Hessian Fusion for Vector Quantization
by: Wang, Zhong, et al.
Published: (2026)
by: Wang, Zhong, et al.
Published: (2026)
Outlier-Aware Training for Low-Bit Quantization of Structural Re-Parameterized Networks
by: Niu, Muqun, et al.
Published: (2024)
by: Niu, Muqun, et al.
Published: (2024)
Fine-tuning Quantized Neural Networks with Zeroth-order Optimization
by: Shang, Sifeng, et al.
Published: (2025)
by: Shang, Sifeng, et al.
Published: (2025)
Transformer-VQ: Linear-Time Transformers via Vector Quantization
by: Lingle, Lucas D.
Published: (2023)
by: Lingle, Lucas D.
Published: (2023)
1-Bit FQT: Pushing the Limit of Fully Quantized Training to 1-bit
by: Gao, Chang, et al.
Published: (2024)
by: Gao, Chang, et al.
Published: (2024)
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference
by: Gafni, Tomer, et al.
Published: (2025)
by: Gafni, Tomer, et al.
Published: (2025)
BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary Codebook
by: Gu, Hao, et al.
Published: (2025)
by: Gu, Hao, et al.
Published: (2025)
EPTQ: Enhanced Post-Training Quantization via Hessian-guided Network-wise Optimization
by: Gordon, Ofir, et al.
Published: (2023)
by: Gordon, Ofir, et al.
Published: (2023)
Quantization Variation: A New Perspective on Training Transformers with Low-Bit Precision
by: Huang, Xijie, et al.
Published: (2023)
by: Huang, Xijie, et al.
Published: (2023)
TWEO: Transformers Without Extreme Outliers Enables FP8 Training And Quantization For Dummies
by: Liang, Guang, et al.
Published: (2025)
by: Liang, Guang, et al.
Published: (2025)
Ovis2.5 Technical Report
by: Lu, Shiyin, et al.
Published: (2025)
by: Lu, Shiyin, et al.
Published: (2025)
When Bits Break Recourse: Counterfactual-Faithful Quantization
by: Yahyati, Chaymae, et al.
Published: (2026)
by: Yahyati, Chaymae, et al.
Published: (2026)
Differentiable, Bit-shifting, and Scalable Quantization without training neural network from scratch
by: Badar, Zia
Published: (2025)
by: Badar, Zia
Published: (2025)
FreeAct: Freeing Activations for LLM Quantization
by: Liu, Xiaohao, et al.
Published: (2026)
by: Liu, Xiaohao, et al.
Published: (2026)
Activation Quantization of Vision Encoders Needs Prefixing Registers
by: Kim, Seunghyeon, et al.
Published: (2025)
by: Kim, Seunghyeon, et al.
Published: (2025)
MimiQ: Low-Bit Data-Free Quantization of Vision Transformers with Encouraging Inter-Head Attention Similarity
by: Choi, Kanghyun, et al.
Published: (2024)
by: Choi, Kanghyun, et al.
Published: (2024)
First-Order Error Matters: Accurate Compensation for Quantized Large Language Models
by: Zheng, Xingyu, et al.
Published: (2025)
by: Zheng, Xingyu, et al.
Published: (2025)
DQA: An Efficient Method for Deep Quantization of Deep Neural Network Activations
by: Hu, Wenhao, et al.
Published: (2024)
by: Hu, Wenhao, et al.
Published: (2024)
UI-Venus-1.5 Technical Report
by: Venus Team, et al.
Published: (2026)
by: Venus Team, et al.
Published: (2026)
Efficient Multi-bit Quantization Network Training via Weight Bias Correction and Bit-wise Coreset Sampling
by: Kim, Jinhee, et al.
Published: (2025)
by: Kim, Jinhee, et al.
Published: (2025)
Post-training Quantization for Text-to-Image Diffusion Models with Progressive Calibration and Activation Relaxing
by: Tang, Siao, et al.
Published: (2023)
by: Tang, Siao, et al.
Published: (2023)
LLM-FP4: 4-Bit Floating-Point Quantized Transformers
by: Liu, Shih-yang, et al.
Published: (2023)
by: Liu, Shih-yang, et al.
Published: (2023)
MLoRQ: Bridging Low-Rank and Quantization for Transformer Compression
by: Gordon, Ofir, et al.
Published: (2025)
by: Gordon, Ofir, et al.
Published: (2025)
MetaMix: Meta-state Precision Searcher for Mixed-precision Activation Quantization
by: Kim, Han-Byul, et al.
Published: (2023)
by: Kim, Han-Byul, et al.
Published: (2023)
Reclaiming Residual Knowledge: A Novel Paradigm to Low-Bit Quantization
by: Luo, Róisín, et al.
Published: (2024)
by: Luo, Róisín, et al.
Published: (2024)
HLQ: Fast and Efficient Backpropagation via Hadamard Low-rank Quantization
by: Kim, Seonggon, et al.
Published: (2024)
by: Kim, Seonggon, et al.
Published: (2024)
Masked Vector Quantization
by: Nguyen, David D., et al.
Published: (2023)
by: Nguyen, David D., et al.
Published: (2023)
Does Vector Quantization Fail in Spatio-Temporal Forecasting? Exploring a Differentiable Sparse Soft-Vector Quantization Approach
by: Chen, Chao, et al.
Published: (2023)
by: Chen, Chao, et al.
Published: (2023)
LUQ: Layerwise Ultra-Low Bit Quantization for Multimodal Large Language Models
by: Bhatnagar, Shubhang, et al.
Published: (2025)
by: Bhatnagar, Shubhang, et al.
Published: (2025)
Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report
by: Lab, Shanghai AI, et al.
Published: (2025)
by: Lab, Shanghai AI, et al.
Published: (2025)
H2OVL-Mississippi Vision Language Models Technical Report
by: Galib, Shaikat, et al.
Published: (2024)
by: Galib, Shaikat, et al.
Published: (2024)
An Empirical Study of World Model Quantization
by: Fu, Zhongqian, et al.
Published: (2026)
by: Fu, Zhongqian, et al.
Published: (2026)
VQ-Style: Disentangling Style and Content in Motion with Residual Quantized Representations
by: Zargarbashi, Fatemeh, et al.
Published: (2026)
by: Zargarbashi, Fatemeh, et al.
Published: (2026)
MPQ-DMv2: Flexible Residual Mixed Precision Quantization for Low-Bit Diffusion Models with Temporal Distillation
by: Feng, Weilun, et al.
Published: (2025)
by: Feng, Weilun, et al.
Published: (2025)
Similar Items
-
Enhancing Post-Training Quantization via Future Activation Awareness
by: Lv, Zheqi, et al.
Published: (2026) -
MARR: Module-Adaptive Residual Reconstruction for Low-Bit Post-Training Quantization
by: Su, Le, et al.
Published: (2026) -
Low-Bit, High-Fidelity: Optimal Transport Quantization for Flow Matching
by: Varam, Dara, et al.
Published: (2025) -
ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization
by: Liu, Zechun, et al.
Published: (2025) -
BayesQ: Uncertainty-Guided Bayesian Quantization
by: Lamaakal, Ismail, et al.
Published: (2025)