FlatQuant: Flatness Matters for LLM Quantization
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Yuxuan, Liu, Ruikang, Bai, Haoli, Bao, Han, Zhao, Kang, Li, Yuening, Hu, Jiaxin, Yu, Xianzhi, Hou, Lu, Yuan, Chun, Jiang, Xin, Liu, Wulong, Yao, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models
by: Liu, Ruikang, et al.
Published: (2025)
by: Liu, Ruikang, et al.
Published: (2025)
IntactKV: Improving Large Language Model Quantization by Keeping Pivot Tokens Intact
by: Liu, Ruikang, et al.
Published: (2024)
by: Liu, Ruikang, et al.
Published: (2024)
QuantClaw: Precision Where It Matters for OpenClaw
by: Zhang, Manyi, et al.
Published: (2026)
by: Zhang, Manyi, et al.
Published: (2026)
A Simple Linear Patch Revives Layer-Pruned Large Language Models
by: Chen, Xinrui, et al.
Published: (2025)
by: Chen, Xinrui, et al.
Published: (2025)
Faster and Better LLMs via Latency-Aware Test-Time Scaling
by: Wang, Zili, et al.
Published: (2025)
by: Wang, Zili, et al.
Published: (2025)
Theory-optimal Quantization Based on Flatness
by: Huang, Xiusheng, et al.
Published: (2026)
by: Huang, Xiusheng, et al.
Published: (2026)
BATQuant: Outlier-resilient MXFP4 Quantization via Learnable Block-wise Optimization
by: Li, Ji-Fu, et al.
Published: (2026)
by: Li, Ji-Fu, et al.
Published: (2026)
What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Study
by: Lv, Keyu, et al.
Published: (2026)
by: Lv, Keyu, et al.
Published: (2026)
Quantization Meets OOD: Generalizable Quantization-aware Training from a Flatness Perspective
by: Jiang, Jiacheng, et al.
Published: (2025)
by: Jiang, Jiacheng, et al.
Published: (2025)
SVDq: 1.25-bit and 410x Key Cache Compression for LLM Attention
by: Yankun, Hong, et al.
Published: (2025)
by: Yankun, Hong, et al.
Published: (2025)
ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
by: Liang, Yesheng, et al.
Published: (2025)
by: Liang, Yesheng, et al.
Published: (2025)
LLMEasyQuant: Scalable Quantization for Parallel and Distributed LLM Inference
by: Liu, Dong, et al.
Published: (2024)
by: Liu, Dong, et al.
Published: (2024)
FedNSAM:Consistency of Local and Global Flatness for Federated Learning
by: Liu, Junkang, et al.
Published: (2026)
by: Liu, Junkang, et al.
Published: (2026)
HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
by: Lin, Haoran, et al.
Published: (2025)
by: Lin, Haoran, et al.
Published: (2025)
Flat Posterior Does Matter For Bayesian Model Averaging
by: Lim, Sungjun, et al.
Published: (2024)
by: Lim, Sungjun, et al.
Published: (2024)
Universal Quantized Berry-Dipole Flat Bands
by: Mo, Qingyang, et al.
Published: (2026)
by: Mo, Qingyang, et al.
Published: (2026)
SliderQuant: Accurate Post-Training Quantization for LLMs
by: Wang, Shigeng, et al.
Published: (2026)
by: Wang, Shigeng, et al.
Published: (2026)
DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization
by: Shao, Yuantian, et al.
Published: (2025)
by: Shao, Yuantian, et al.
Published: (2025)
FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization
by: Liu, Fangxin, et al.
Published: (2025)
by: Liu, Fangxin, et al.
Published: (2025)
MixPE: Quantization and Hardware Co-design for Efficient LLM Inference
by: Zhang, Yu, et al.
Published: (2024)
by: Zhang, Yu, et al.
Published: (2024)
FlatFormer: A Flat Transformer Knowledge Tracing Model Based on Cognitive Bias Injection
by: Xia, Xiao-li, et al.
Published: (2025)
by: Xia, Xiao-li, et al.
Published: (2025)
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
by: Xu, Zukang, et al.
Published: (2025)
by: Xu, Zukang, et al.
Published: (2025)
Benchmarking Post-Training Quantization of Large Language Models under Microscaling Floating Point Formats
by: Zhang, Manyi, et al.
Published: (2026)
by: Zhang, Manyi, et al.
Published: (2026)
EasyQuant: An Efficient Data-free Quantization Algorithm for LLMs
by: Tang, Hanlin, et al.
Published: (2024)
by: Tang, Hanlin, et al.
Published: (2024)
On Two Dimensional Flat Hessian Potentials
by: Liu, Hanwen
Published: (2025)
by: Liu, Hanwen
Published: (2025)
A Deformation Quantization for Non-Flat Spacetimes and Applications to QFT
by: Much, Albert
Published: (2021)
by: Much, Albert
Published: (2021)
E$^3$-Pruner: Towards Efficient, Economical, and Effective Layer Pruning for Large Language Models
by: Yuan, Tao, et al.
Published: (2025)
by: Yuan, Tao, et al.
Published: (2025)
MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM
by: Wang, Dongwei, et al.
Published: (2026)
by: Wang, Dongwei, et al.
Published: (2026)
DynaQuant: Dynamic Mixed-Precision Quantization for Learned Image Compression
by: Bao, Youneng, et al.
Published: (2025)
by: Bao, Youneng, et al.
Published: (2025)
Giant Magneto-Optical Effects in Two-Dimensional Flat-Band Antiferromagnets
by: Yang, Ping, et al.
Published: (2025)
by: Yang, Ping, et al.
Published: (2025)
Characterizations of Griffiths Positivity, Pluriharmonicity and Flatness
by: Liu, Zhuo, et al.
Published: (2022)
by: Liu, Zhuo, et al.
Published: (2022)
StableQuant: Layer Adaptive Post-Training Quantization for Speech Foundation Models
by: Hong, Yeona, et al.
Published: (2025)
by: Hong, Yeona, et al.
Published: (2025)
Identifying Research Trends, Active Research Areas and Pivotal Publications with Co-Citation Analysis in CiteSpace: A Case Study with Active Matter
by: Yuening Zhang
Published: (2025)
by: Yuening Zhang
Published: (2025)
Levi Flat Structures via Structure Sheaves: Differential Complexes, Convexity, and Global Solvability
by: Ji, Qingchun, et al.
Published: (2025)
by: Ji, Qingchun, et al.
Published: (2025)
Symmetry-Based Real-Space Framework for Realizing Flat Bands and Unveiling Nodal-Line Touchings
by: Liu, Rui-Heng, et al.
Published: (2024)
by: Liu, Rui-Heng, et al.
Published: (2024)
AttentionPredictor: Temporal Patterns Matter for KV Cache Compression
by: Yang, Qingyue, et al.
Published: (2025)
by: Yang, Qingyue, et al.
Published: (2025)
Flat Band Josephson Junctions with Quantum Metric
by: Li, Zhong C. F., et al.
Published: (2024)
by: Li, Zhong C. F., et al.
Published: (2024)
GS-Quant: Granular Semantic and Generative Structural Quantization for Knowledge Graph Completion
by: Xie, Qizhuo, et al.
Published: (2026)
by: Xie, Qizhuo, et al.
Published: (2026)
DilateQuant: Accurate and Efficient Diffusion Quantization via Weight Dilation
by: Liu, Xuewen, et al.
Published: (2024)
by: Liu, Xuewen, et al.
Published: (2024)
One QuantLLM for ALL: Fine-tuning Quantized LLMs Once for Efficient Deployments
by: Yi, Ke, et al.
Published: (2024)
by: Yi, Ke, et al.
Published: (2024)
Similar Items
-
Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models
by: Liu, Ruikang, et al.
Published: (2025) -
IntactKV: Improving Large Language Model Quantization by Keeping Pivot Tokens Intact
by: Liu, Ruikang, et al.
Published: (2024) -
QuantClaw: Precision Where It Matters for OpenClaw
by: Zhang, Manyi, et al.
Published: (2026) -
A Simple Linear Patch Revives Layer-Pruned Large Language Models
by: Chen, Xinrui, et al.
Published: (2025) -
Faster and Better LLMs via Latency-Aware Test-Time Scaling
by: Wang, Zili, et al.
Published: (2025)