Achieving binary weight and activation for LLMs using Post-Training Quantization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Siqing, Wang, Chuang, Wang, Ruiqi, Yang, Yi, Zhang, Xu-Yao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913918710972416
author Song, Siqing
Wang, Chuang
Wang, Ruiqi
Yang, Yi
Zhang, Xu-Yao
author_facet Song, Siqing
Wang, Chuang
Wang, Ruiqi
Yang, Yi
Zhang, Xu-Yao
contents Quantizing large language models (LLMs) to 1-bit precision significantly reduces computational costs, but existing quantization techniques suffer from noticeable performance degradation when using weight and activation precisions below 4 bits (W4A4). In this paper, we propose a post-training quantization framework with W(1+1)A(1*4) configuration, where weights are quantized to 1 bit with an additional 1 bit for fine-grain grouping and activations are quantized to 1 bit with a 4-fold increase in the number of channels. For weight quantization, we propose utilizing Hessian-aware fine-grained grouping along with an EM-based quantization scheme. For activation quantization, we decompose INT4-quantized activations into a 4 * INT1 format equivalently and simultaneously smooth the scaling factors based on quantization errors, which further reduces the quantization errors in activations. Our method surpasses state-of-the-art (SOTA) LLM quantization baselines on W2A4 across multiple tasks, pushing the boundaries of existing LLM quantization methods toward fully binarized models. Code is available at https://github.com/JimmyCrave/LLM-PTQ-binarization.
format Preprint
id arxiv_https___arxiv_org_abs_2504_05352
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Achieving binary weight and activation for LLMs using Post-Training Quantization
Song, Siqing
Wang, Chuang
Wang, Ruiqi
Yang, Yi
Zhang, Xu-Yao
Machine Learning
Artificial Intelligence
Quantizing large language models (LLMs) to 1-bit precision significantly reduces computational costs, but existing quantization techniques suffer from noticeable performance degradation when using weight and activation precisions below 4 bits (W4A4). In this paper, we propose a post-training quantization framework with W(1+1)A(1*4) configuration, where weights are quantized to 1 bit with an additional 1 bit for fine-grain grouping and activations are quantized to 1 bit with a 4-fold increase in the number of channels. For weight quantization, we propose utilizing Hessian-aware fine-grained grouping along with an EM-based quantization scheme. For activation quantization, we decompose INT4-quantized activations into a 4 * INT1 format equivalently and simultaneously smooth the scaling factors based on quantization errors, which further reduces the quantization errors in activations. Our method surpasses state-of-the-art (SOTA) LLM quantization baselines on W2A4 across multiple tasks, pushing the boundaries of existing LLM quantization methods toward fully binarized models. Code is available at https://github.com/JimmyCrave/LLM-PTQ-binarization.
title Achieving binary weight and activation for LLMs using Post-Training Quantization
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2504.05352