HeRo-Q: A General Framework for Stable Low Bit Quantization via Hessian Conditioning
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Jinhao Zhang Yunquan, yan, Zicheng, Zhang, Boyang, Sun, Jun, Cheng, Daning |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CALM: A CKA-Guided Adaptive Layer-Wise Modularization Framework for LLM Quantization
by: Zhang, Jinhao, et al.
Published: (2025)
by: Zhang, Jinhao, et al.
Published: (2025)
MoQE: Improve Quantization Model performance via Mixture of Quantization Experts
by: Zhang, Jinhao, et al.
Published: (2025)
by: Zhang, Jinhao, et al.
Published: (2025)
FP=xINT:Representing Neural Networks via Low-Bit Series Basis Functions
by: Zhang, Boyang, et al.
Published: (2024)
by: Zhang, Boyang, et al.
Published: (2024)
Compression for Better: A General and Stable Lossless Compression Framework
by: Zhang, Boyang, et al.
Published: (2024)
by: Zhang, Boyang, et al.
Published: (2024)
MoE-DisCo:Low Economy Cost Training Mixture-of-Experts Models
by: Ye, Xin, et al.
Published: (2026)
by: Ye, Xin, et al.
Published: (2026)
Lossless Model Compression via Joint Low-Rank Factorization Optimization
by: Zhang, Boyang, et al.
Published: (2024)
by: Zhang, Boyang, et al.
Published: (2024)
A General Error-Theoretical Analysis Framework for Constructing Compression Strategies
by: Zhang, Boyang, et al.
Published: (2025)
by: Zhang, Boyang, et al.
Published: (2025)
A Qualitative Test-Risk Mechanism for Scaling Behavior in Normalized Residual Networks
by: Cheng, Daning, et al.
Published: (2026)
by: Cheng, Daning, et al.
Published: (2026)
Exploiting Block Coordinate Descent for Cost-Effective LLM Model Training
by: Liu, Zeyu, et al.
Published: (2025)
by: Liu, Zeyu, et al.
Published: (2025)
Rethinking Parameter Sharing as Graph Coloring for Structured Compression
by: Zhang, Boyang, et al.
Published: (2025)
by: Zhang, Boyang, et al.
Published: (2025)
HESTIA: A Hessian-Guided Differentiable Quantization-Aware Training Framework for Extremely Low-Bit LLMs
by: Wang, Guoan, et al.
Published: (2026)
by: Wang, Guoan, et al.
Published: (2026)
Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models
by: Zhong, Yi, et al.
Published: (2026)
by: Zhong, Yi, et al.
Published: (2026)
CaRoBio: 3D Cable Routing with a Bio-inspired Gripper Fingernail
by: Zuo, Jiahui, et al.
Published: (2025)
by: Zuo, Jiahui, et al.
Published: (2025)
Q$^2$: Quantization-Aware Gradient Balancing and Attention Alignment for Low-Bit Quantization
by: Wang, Zhaoyang, et al.
Published: (2025)
by: Wang, Zhaoyang, et al.
Published: (2025)
HeRo: Adaptive Orchestration of Agentic RAG on Heterogeneous Mobile SoC
by: Li, Maoliang, et al.
Published: (2026)
by: Li, Maoliang, et al.
Published: (2026)
LittleBit: Ultra Low-Bit Quantization via Latent Factorization
by: Lee, Banseok, et al.
Published: (2025)
by: Lee, Banseok, et al.
Published: (2025)
Learning Grouped Lattice Vector Quantizers for Low-Bit LLM Compression
by: Zhang, Xi, et al.
Published: (2025)
by: Zhang, Xi, et al.
Published: (2025)
Why Do Some Inputs Break Low-Bit LLM Quantization?
by: Chang, Ting-Yun, et al.
Published: (2025)
by: Chang, Ting-Yun, et al.
Published: (2025)
LFQ: Logit-aware Final-block Quantization for Boosting the Generation Quality of Low-Bit Quantized LLMs
by: Lee, Jung Hyun, et al.
Published: (2026)
by: Lee, Jung Hyun, et al.
Published: (2026)
Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
by: Zhang, Tuo, et al.
Published: (2025)
by: Zhang, Tuo, et al.
Published: (2025)
Verification of Bit-Flip Attacks against Quantized Neural Networks
by: Zhang, Yedi, et al.
Published: (2025)
by: Zhang, Yedi, et al.
Published: (2025)
Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM Deployment
by: Lee, Deokjae, et al.
Published: (2025)
by: Lee, Deokjae, et al.
Published: (2025)
SignRoundV2: Toward Closing the Performance Gap in Extremely Low-Bit Post-Training Quantization for LLMs
by: Cheng, Wenhua, et al.
Published: (2025)
by: Cheng, Wenhua, et al.
Published: (2025)
Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
by: Zhang, Tianao, et al.
Published: (2025)
by: Zhang, Tianao, et al.
Published: (2025)
Can the capability of Large Language Models be described by human ability? A Meta Study
by: Zan, Mingrui, et al.
Published: (2025)
by: Zan, Mingrui, et al.
Published: (2025)
Attn-QAT: 4-Bit Attention With Quantization-Aware Training
by: Zhang, Peiyuan, et al.
Published: (2026)
by: Zhang, Peiyuan, et al.
Published: (2026)
MimiQ: Low-Bit Data-Free Quantization of Vision Transformers with Encouraging Inter-Head Attention Similarity
by: Choi, Kanghyun, et al.
Published: (2024)
by: Choi, Kanghyun, et al.
Published: (2024)
PTQ1.61: Push the Real Limit of Extremely Low-Bit Post-Training Quantization Methods for Large Language Models
by: Zhao, Jiaqi, et al.
Published: (2025)
by: Zhao, Jiaqi, et al.
Published: (2025)
Quantization Meets Reasoning: Exploring LLM Low-Bit Quantization Degradation for Mathematical Reasoning
by: Li, Zhen, et al.
Published: (2025)
by: Li, Zhen, et al.
Published: (2025)
EliteKV: Scalable KV Cache Compression via RoPE Frequency Selection and Joint Low-Rank Projection
by: Zhou, Yuhao, et al.
Published: (2025)
by: Zhou, Yuhao, et al.
Published: (2025)
Evolution Strategy-Based Calibration for Low-Bit Quantization of Speech Models
by: Rakotoarivony, Lucas
Published: (2026)
by: Rakotoarivony, Lucas
Published: (2026)
R2Q: Towards Robust 2-Bit Large Language Models via Residual Refinement Quantization
by: Chen, Jiayi, et al.
Published: (2025)
by: Chen, Jiayi, et al.
Published: (2025)
Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision
by: Wu, Haoning, et al.
Published: (2023)
by: Wu, Haoning, et al.
Published: (2023)
Q-ROAR: Outlier-Aware Rescaling for RoPE Position Interpolation in Quantized Long-Context LLMs
by: Qiao, Ye, et al.
Published: (2025)
by: Qiao, Ye, et al.
Published: (2025)
InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization
by: Li, Ke, et al.
Published: (2026)
by: Li, Ke, et al.
Published: (2026)
SplitQuant: Layer Splitting for Low-Bit Neural Network Quantization
by: Song, Jaewoo, et al.
Published: (2025)
by: Song, Jaewoo, et al.
Published: (2025)
Quantization Variation: A New Perspective on Training Transformers with Low-Bit Precision
by: Huang, Xijie, et al.
Published: (2023)
by: Huang, Xijie, et al.
Published: (2023)
RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache
by: Zhang, Junkai, et al.
Published: (2026)
by: Zhang, Junkai, et al.
Published: (2026)
ECQ$^{\text{x}}$: Explainability-Driven Quantization for Low-Bit and Sparse DNNs
by: Becking, Daniel, et al.
Published: (2021)
by: Becking, Daniel, et al.
Published: (2021)
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
by: Yan, Xianglong, et al.
Published: (2026)
by: Yan, Xianglong, et al.
Published: (2026)
Similar Items
-
CALM: A CKA-Guided Adaptive Layer-Wise Modularization Framework for LLM Quantization
by: Zhang, Jinhao, et al.
Published: (2025) -
MoQE: Improve Quantization Model performance via Mixture of Quantization Experts
by: Zhang, Jinhao, et al.
Published: (2025) -
FP=xINT:Representing Neural Networks via Low-Bit Series Basis Functions
by: Zhang, Boyang, et al.
Published: (2024) -
Compression for Better: A General and Stable Lossless Compression Framework
by: Zhang, Boyang, et al.
Published: (2024) -
MoE-DisCo:Low Economy Cost Training Mixture-of-Experts Models
by: Ye, Xin, et al.
Published: (2026)