RaBiT: Residual-Aware Binarization Training for Accurate and Efficient LLMs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: You, Youngcheon, Lee, Banseok, Choi, Minseop, Kim, Seonyoung, Chong, Hyochan, Kim, Changdong, Kim, Youngmin, Kim, Dongkyu
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910224226451456
author You, Youngcheon
Lee, Banseok
Choi, Minseop
Kim, Seonyoung
Chong, Hyochan
Kim, Changdong
Kim, Youngmin
Kim, Dongkyu
author_facet You, Youngcheon
Lee, Banseok
Choi, Minseop
Kim, Seonyoung
Chong, Hyochan
Kim, Changdong
Kim, Youngmin
Kim, Dongkyu
contents Efficient deployment of large language models (LLMs) requires extreme quantization, forcing a critical trade-off between low-bit efficiency and performance. Residual binarization enables hardware-friendly, matmul-free inference by stacking binary ($\pm$1) layers, but is plagued by pathological feature co-adaptation. We identify a key failure mode, which we term inter-path adaptation: during quantization-aware training (QAT), parallel residual binary paths learn redundant features, degrading the error-compensation structure and limiting the expressive capacity of the model. While prior work relies on heuristic workarounds (e.g., path freezing) that constrain the solution space, we propose RaBiT, a novel quantization framework that resolves co-adaptation by algorithmically enforcing a residual hierarchy. Its core mechanism sequentially derives each binary path from a single shared full-precision weight, which ensures that every path corrects the error of the preceding one. This process is stabilized by a robust initialization that prioritizes functional preservation over mere weight approximation. RaBiT redefines the 2-bit accuracy-efficiency frontier: it achieves state-of-the-art performance, rivals even hardware-intensive Vector Quantization (VQ) methods, and delivers a $4.49\times$ inference speed-up over full-precision models on an RTX 4090.
format Preprint
id arxiv_https___arxiv_org_abs_2602_05367
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RaBiT: Residual-Aware Binarization Training for Accurate and Efficient LLMs
You, Youngcheon
Lee, Banseok
Choi, Minseop
Kim, Seonyoung
Chong, Hyochan
Kim, Changdong
Kim, Youngmin
Kim, Dongkyu
Artificial Intelligence
Efficient deployment of large language models (LLMs) requires extreme quantization, forcing a critical trade-off between low-bit efficiency and performance. Residual binarization enables hardware-friendly, matmul-free inference by stacking binary ($\pm$1) layers, but is plagued by pathological feature co-adaptation. We identify a key failure mode, which we term inter-path adaptation: during quantization-aware training (QAT), parallel residual binary paths learn redundant features, degrading the error-compensation structure and limiting the expressive capacity of the model. While prior work relies on heuristic workarounds (e.g., path freezing) that constrain the solution space, we propose RaBiT, a novel quantization framework that resolves co-adaptation by algorithmically enforcing a residual hierarchy. Its core mechanism sequentially derives each binary path from a single shared full-precision weight, which ensures that every path corrects the error of the preceding one. This process is stabilized by a robust initialization that prioritizes functional preservation over mere weight approximation. RaBiT redefines the 2-bit accuracy-efficiency frontier: it achieves state-of-the-art performance, rivals even hardware-intensive Vector Quantization (VQ) methods, and delivers a $4.49\times$ inference speed-up over full-precision models on an RTX 4090.
title RaBiT: Residual-Aware Binarization Training for Accurate and Efficient LLMs
topic Artificial Intelligence
url https://arxiv.org/abs/2602.05367