Achieving binary weight and activation for LLMs using Post-Training Quantization
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Siqing, Wang, Chuang, Wang, Ruiqi, Yang, Yi, Zhang, Xu-Yao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LBLLM: Lightweight Binarization of Large Language Models via Three-Stage Distillation
by: Song, Siqing, et al.
Published: (2026)
by: Song, Siqing, et al.
Published: (2026)
BWLA: Breaking the Barrier of W1AX Post-Training Quantization for LLMs
by: Zhao, Zhixiong, et al.
Published: (2026)
by: Zhao, Zhixiong, et al.
Published: (2026)
Benchmarking Post-Training Quantization in LLMs: Comprehensive Taxonomy, Unified Evaluation, and Comparative Analysis
by: Zhao, Jiaqi, et al.
Published: (2025)
by: Zhao, Jiaqi, et al.
Published: (2025)
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
by: Yan, Xianglong, et al.
Published: (2026)
by: Yan, Xianglong, et al.
Published: (2026)
MagR: Weight Magnitude Reduction for Enhancing Post-Training Quantization
by: Zhang, Aozhong, et al.
Published: (2024)
by: Zhang, Aozhong, et al.
Published: (2024)
DAQ: Density-Aware Post-Training Weight-Only Quantization For LLMs
by: Luo, Yingsong, et al.
Published: (2024)
by: Luo, Yingsong, et al.
Published: (2024)
Post Training Quantization of Large Language Models with Microscaling Formats
by: Sharify, Sayeh, et al.
Published: (2024)
by: Sharify, Sayeh, et al.
Published: (2024)
BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
by: Huang, Wei, et al.
Published: (2024)
by: Huang, Wei, et al.
Published: (2024)
PTQTP: Post-Training Quantization to Trit-Planes for Large Language Models
by: Xiao, He, et al.
Published: (2025)
by: Xiao, He, et al.
Published: (2025)
Activation Sensitivity as a Unifying Principle for Post-Training Quantization
by: Xu, Bruce Changlong
Published: (2026)
by: Xu, Bruce Changlong
Published: (2026)
Beacon: Post-Training Quantization with Integrated Grid Selection
by: Zhang, Shihao, et al.
Published: (2025)
by: Zhang, Shihao, et al.
Published: (2025)
CrossQuant: A Post-Training Quantization Method with Smaller Quantization Kernel for Precise Large Language Model Compression
by: Liu, Wenyuan, et al.
Published: (2024)
by: Liu, Wenyuan, et al.
Published: (2024)
DAQ: Delta-Aware Quantization for Post-Training LLM Weight Compression
by: Yu, Xiaoming, et al.
Published: (2026)
by: Yu, Xiaoming, et al.
Published: (2026)
MoBiE: Efficient Inference of Mixture of Binary Experts under Post-Training Quantization
by: Zhao, Zhixiong, et al.
Published: (2026)
by: Zhao, Zhixiong, et al.
Published: (2026)
HESTIA: A Hessian-Guided Differentiable Quantization-Aware Training Framework for Extremely Low-Bit LLMs
by: Wang, Guoan, et al.
Published: (2026)
by: Wang, Guoan, et al.
Published: (2026)
Astro: Activation-guided Structured Regularization for Outlier-Robust LLM Post-Training Quantization
by: Chen, Xi, et al.
Published: (2026)
by: Chen, Xi, et al.
Published: (2026)
Pushing the Limits of Block Rotations in Post-Training Quantization
by: Sanjeet, Sai, et al.
Published: (2026)
by: Sanjeet, Sai, et al.
Published: (2026)
Exploring Layer-wise Information Effectiveness for Post-Training Quantization in Small Language Models
by: Xiao, He, et al.
Published: (2025)
by: Xiao, He, et al.
Published: (2025)
A Quantized VAE-MLP Botnet Detection Model: A Systematic Evaluation of Quantization-Aware Training and Post-Training Quantization Strategies
by: Wasswa, Hassan, et al.
Published: (2025)
by: Wasswa, Hassan, et al.
Published: (2025)
PTQ1.61: Push the Real Limit of Extremely Low-Bit Post-Training Quantization Methods for Large Language Models
by: Zhao, Jiaqi, et al.
Published: (2025)
by: Zhao, Jiaqi, et al.
Published: (2025)
AMS-QUANT: Adaptive Mantissa Sharing for Floating-point Quantization
by: Lv, Mengtao, et al.
Published: (2025)
by: Lv, Mengtao, et al.
Published: (2025)
RaanA: A Fast, Flexible, and Data-Efficient Post-Training Quantization Algorithm
by: Yang, Yongyi, et al.
Published: (2025)
by: Yang, Yongyi, et al.
Published: (2025)
End-to-End On-Device Quantization-Aware Training for LLMs at Inference Cost
by: Tan, Qitao, et al.
Published: (2025)
by: Tan, Qitao, et al.
Published: (2025)
Lightweight and Post-Training Structured Pruning for On-Device Large Lanaguage Models
by: Xu, Zihuai, et al.
Published: (2025)
by: Xu, Zihuai, et al.
Published: (2025)
Qronos: Correcting the Past by Shaping the Future... in Post-Training Quantization
by: Zhang, Shihao, et al.
Published: (2025)
by: Zhang, Shihao, et al.
Published: (2025)
Post-Training Quantization of OpenPangu Models for Efficient Deployment on Atlas A2
by: Luo, Yilun, et al.
Published: (2025)
by: Luo, Yilun, et al.
Published: (2025)
Towards Next-Level Post-Training Quantization of Hyper-Scale Transformers
by: Kim, Junhan, et al.
Published: (2024)
by: Kim, Junhan, et al.
Published: (2024)
Fira: Can We Achieve Full-rank Training of LLMs Under Low-rank Constraint?
by: Chen, Xi, et al.
Published: (2024)
by: Chen, Xi, et al.
Published: (2024)
Improving Quantization with Post-Training Model Expansion
by: Franco, Giuseppe, et al.
Published: (2025)
by: Franco, Giuseppe, et al.
Published: (2025)
FP4 All the Way: Fully Quantized Training of LLMs
by: Chmiel, Brian, et al.
Published: (2025)
by: Chmiel, Brian, et al.
Published: (2025)
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
by: Liu, Yixin, et al.
Published: (2026)
by: Liu, Yixin, et al.
Published: (2026)
Accumulator-Aware Post-Training Quantization for Large Language Models
by: Colbert, Ian, et al.
Published: (2024)
by: Colbert, Ian, et al.
Published: (2024)
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding
by: Zhoubian, Sining, et al.
Published: (2025)
by: Zhoubian, Sining, et al.
Published: (2025)
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
by: Xiao, Guangxuan, et al.
Published: (2022)
by: Xiao, Guangxuan, et al.
Published: (2022)
Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
by: Zhang, Tianao, et al.
Published: (2025)
by: Zhang, Tianao, et al.
Published: (2025)
Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
by: Zhang, Tuo, et al.
Published: (2025)
by: Zhang, Tuo, et al.
Published: (2025)
Quamba: A Post-Training Quantization Recipe for Selective State Space Models
by: Chiang, Hung-Yueh, et al.
Published: (2024)
by: Chiang, Hung-Yueh, et al.
Published: (2024)
LRQ: Optimizing Post-Training Quantization for Large Language Models by Learning Low-Rank Weight-Scaling Matrices
by: Lee, Jung Hyun, et al.
Published: (2024)
by: Lee, Jung Hyun, et al.
Published: (2024)
Athena: Efficient Block-Wise Post-Training Quantization for Large Language Models Using Second-Order Matrix Derivative Information
by: Wang, Yanshu, et al.
Published: (2024)
by: Wang, Yanshu, et al.
Published: (2024)
FIMA-Q: Post-Training Quantization for Vision Transformers by Fisher Information Matrix Approximation
by: Wu, Zhuguanyu, et al.
Published: (2025)
by: Wu, Zhuguanyu, et al.
Published: (2025)
Similar Items
-
LBLLM: Lightweight Binarization of Large Language Models via Three-Stage Distillation
by: Song, Siqing, et al.
Published: (2026) -
BWLA: Breaking the Barrier of W1AX Post-Training Quantization for LLMs
by: Zhao, Zhixiong, et al.
Published: (2026) -
Benchmarking Post-Training Quantization in LLMs: Comprehensive Taxonomy, Unified Evaluation, and Comparative Analysis
by: Zhao, Jiaqi, et al.
Published: (2025) -
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
by: Yan, Xianglong, et al.
Published: (2026) -
MagR: Weight Magnitude Reduction for Enhancing Post-Training Quantization
by: Zhang, Aozhong, et al.
Published: (2024)