MagR: Weight Magnitude Reduction for Enhancing Post-Training Quantization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Aozhong, Wang, Naigang, Deng, Yanxia, Li, Xin, Yang, Zi, Yin, Penghang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CLoQ: Enhancing Fine-Tuning of Quantized LLMs via Calibrated LoRA Initialization
von: Deng, Yanxia, et al.
Veröffentlicht: (2025)
von: Deng, Yanxia, et al.
Veröffentlicht: (2025)
DiaBlo: Diagonal Blocks Are Sufficient For Finetuning
von: Gurses, Selcuk, et al.
Veröffentlicht: (2025)
von: Gurses, Selcuk, et al.
Veröffentlicht: (2025)
COMQ: A Backpropagation-Free Algorithm for Post-Training Quantization
von: Zhang, Aozhong, et al.
Veröffentlicht: (2024)
von: Zhang, Aozhong, et al.
Veröffentlicht: (2024)
Frayed RoPE and Long Inputs: A Geometric Perspective
von: Wertheimer, Davis, et al.
Veröffentlicht: (2026)
von: Wertheimer, Davis, et al.
Veröffentlicht: (2026)
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
von: Yan, Xianglong, et al.
Veröffentlicht: (2026)
von: Yan, Xianglong, et al.
Veröffentlicht: (2026)
DAQ: Delta-Aware Quantization for Post-Training LLM Weight Compression
von: Yu, Xiaoming, et al.
Veröffentlicht: (2026)
von: Yu, Xiaoming, et al.
Veröffentlicht: (2026)
DAQ: Density-Aware Post-Training Weight-Only Quantization For LLMs
von: Luo, Yingsong, et al.
Veröffentlicht: (2024)
von: Luo, Yingsong, et al.
Veröffentlicht: (2024)
LRQ: Optimizing Post-Training Quantization for Large Language Models by Learning Low-Rank Weight-Scaling Matrices
von: Lee, Jung Hyun, et al.
Veröffentlicht: (2024)
von: Lee, Jung Hyun, et al.
Veröffentlicht: (2024)
Post Training Quantization of Large Language Models with Microscaling Formats
von: Sharify, Sayeh, et al.
Veröffentlicht: (2024)
von: Sharify, Sayeh, et al.
Veröffentlicht: (2024)
DNN Memory Footprint Reduction via Post-Training Intra-Layer Multi-Precision Quantization
von: Ghavami, Behnam, et al.
Veröffentlicht: (2024)
von: Ghavami, Behnam, et al.
Veröffentlicht: (2024)
Beacon: Post-Training Quantization with Integrated Grid Selection
von: Zhang, Shihao, et al.
Veröffentlicht: (2025)
von: Zhang, Shihao, et al.
Veröffentlicht: (2025)
PTQTP: Post-Training Quantization to Trit-Planes for Large Language Models
von: Xiao, He, et al.
Veröffentlicht: (2025)
von: Xiao, He, et al.
Veröffentlicht: (2025)
Achieving binary weight and activation for LLMs using Post-Training Quantization
von: Song, Siqing, et al.
Veröffentlicht: (2025)
von: Song, Siqing, et al.
Veröffentlicht: (2025)
Pushing the Limits of Block Rotations in Post-Training Quantization
von: Sanjeet, Sai, et al.
Veröffentlicht: (2026)
von: Sanjeet, Sai, et al.
Veröffentlicht: (2026)
BWLA: Breaking the Barrier of W1AX Post-Training Quantization for LLMs
von: Zhao, Zhixiong, et al.
Veröffentlicht: (2026)
von: Zhao, Zhixiong, et al.
Veröffentlicht: (2026)
A Quantized VAE-MLP Botnet Detection Model: A Systematic Evaluation of Quantization-Aware Training and Post-Training Quantization Strategies
von: Wasswa, Hassan, et al.
Veröffentlicht: (2025)
von: Wasswa, Hassan, et al.
Veröffentlicht: (2025)
OptRot: Mitigating Weight Outliers via Data-Free Rotations for Post-Training Quantization
von: Gadhikar, Advait, et al.
Veröffentlicht: (2025)
von: Gadhikar, Advait, et al.
Veröffentlicht: (2025)
Efficient Quantization of Mixture-of-Experts with Theoretical Generalization Guarantees
von: Chowdhury, Mohammed Nowaz Rabbani, et al.
Veröffentlicht: (2026)
von: Chowdhury, Mohammed Nowaz Rabbani, et al.
Veröffentlicht: (2026)
RaanA: A Fast, Flexible, and Data-Efficient Post-Training Quantization Algorithm
von: Yang, Yongyi, et al.
Veröffentlicht: (2025)
von: Yang, Yongyi, et al.
Veröffentlicht: (2025)
Activation Sensitivity as a Unifying Principle for Post-Training Quantization
von: Xu, Bruce Changlong
Veröffentlicht: (2026)
von: Xu, Bruce Changlong
Veröffentlicht: (2026)
Astro: Activation-guided Structured Regularization for Outlier-Robust LLM Post-Training Quantization
von: Chen, Xi, et al.
Veröffentlicht: (2026)
von: Chen, Xi, et al.
Veröffentlicht: (2026)
CrossQuant: A Post-Training Quantization Method with Smaller Quantization Kernel for Precise Large Language Model Compression
von: Liu, Wenyuan, et al.
Veröffentlicht: (2024)
von: Liu, Wenyuan, et al.
Veröffentlicht: (2024)
Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
von: Zhang, Tianao, et al.
Veröffentlicht: (2025)
von: Zhang, Tianao, et al.
Veröffentlicht: (2025)
Sparse Weight Averaging with Multiple Particles for Iterative Magnitude Pruning
von: Choi, Moonseok, et al.
Veröffentlicht: (2023)
von: Choi, Moonseok, et al.
Veröffentlicht: (2023)
Benchmarking Post-Training Quantization in LLMs: Comprehensive Taxonomy, Unified Evaluation, and Comparative Analysis
von: Zhao, Jiaqi, et al.
Veröffentlicht: (2025)
von: Zhao, Jiaqi, et al.
Veröffentlicht: (2025)
Qronos: Correcting the Past by Shaping the Future... in Post-Training Quantization
von: Zhang, Shihao, et al.
Veröffentlicht: (2025)
von: Zhang, Shihao, et al.
Veröffentlicht: (2025)
BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
von: Huang, Wei, et al.
Veröffentlicht: (2024)
von: Huang, Wei, et al.
Veröffentlicht: (2024)
MoBiE: Efficient Inference of Mixture of Binary Experts under Post-Training Quantization
von: Zhao, Zhixiong, et al.
Veröffentlicht: (2026)
von: Zhao, Zhixiong, et al.
Veröffentlicht: (2026)
Post-Training Quantization of OpenPangu Models for Efficient Deployment on Atlas A2
von: Luo, Yilun, et al.
Veröffentlicht: (2025)
von: Luo, Yilun, et al.
Veröffentlicht: (2025)
Mag-Mamba: Modeling Coupled spatiotemporal Asymmetry for POI Recommendation
von: Li, Zhuoxuan, et al.
Veröffentlicht: (2026)
von: Li, Zhuoxuan, et al.
Veröffentlicht: (2026)
Towards Next-Level Post-Training Quantization of Hyper-Scale Transformers
von: Kim, Junhan, et al.
Veröffentlicht: (2024)
von: Kim, Junhan, et al.
Veröffentlicht: (2024)
Exploring Layer-wise Information Effectiveness for Post-Training Quantization in Small Language Models
von: Xiao, He, et al.
Veröffentlicht: (2025)
von: Xiao, He, et al.
Veröffentlicht: (2025)
Improving Quantization with Post-Training Model Expansion
von: Franco, Giuseppe, et al.
Veröffentlicht: (2025)
von: Franco, Giuseppe, et al.
Veröffentlicht: (2025)
Accumulator-Aware Post-Training Quantization for Large Language Models
von: Colbert, Ian, et al.
Veröffentlicht: (2024)
von: Colbert, Ian, et al.
Veröffentlicht: (2024)
Quamba: A Post-Training Quantization Recipe for Selective State Space Models
von: Chiang, Hung-Yueh, et al.
Veröffentlicht: (2024)
von: Chiang, Hung-Yueh, et al.
Veröffentlicht: (2024)
FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization
von: Lee, Jung Hyun, et al.
Veröffentlicht: (2023)
von: Lee, Jung Hyun, et al.
Veröffentlicht: (2023)
ECO: Quantized Training without Full-Precision Master Weights
von: Nikdan, Mahdi, et al.
Veröffentlicht: (2026)
von: Nikdan, Mahdi, et al.
Veröffentlicht: (2026)
The Primacy of Magnitude in Low-Rank Adaptation
von: Zhang, Zicheng, et al.
Veröffentlicht: (2025)
von: Zhang, Zicheng, et al.
Veröffentlicht: (2025)
PTQ1.61: Push the Real Limit of Extremely Low-Bit Post-Training Quantization Methods for Large Language Models
von: Zhao, Jiaqi, et al.
Veröffentlicht: (2025)
von: Zhao, Jiaqi, et al.
Veröffentlicht: (2025)
QuantMoE-Bench: Examining Post-Training Quantization for Mixture-of-Experts
von: Li, Pingzhi, et al.
Veröffentlicht: (2024)
von: Li, Pingzhi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
CLoQ: Enhancing Fine-Tuning of Quantized LLMs via Calibrated LoRA Initialization
von: Deng, Yanxia, et al.
Veröffentlicht: (2025) -
DiaBlo: Diagonal Blocks Are Sufficient For Finetuning
von: Gurses, Selcuk, et al.
Veröffentlicht: (2025) -
COMQ: A Backpropagation-Free Algorithm for Post-Training Quantization
von: Zhang, Aozhong, et al.
Veröffentlicht: (2024) -
Frayed RoPE and Long Inputs: A Geometric Perspective
von: Wertheimer, Davis, et al.
Veröffentlicht: (2026) -
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
von: Yan, Xianglong, et al.
Veröffentlicht: (2026)