Mixture of Scales: Memory-Efficient Token-Adaptive Binarization for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jo, Dongwon, Kim, Taesu, Kim, Yulhwa, Kim, Jae-Joon
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912094898618368
author Jo, Dongwon
Kim, Taesu
Kim, Yulhwa
Kim, Jae-Joon
author_facet Jo, Dongwon
Kim, Taesu
Kim, Yulhwa
Kim, Jae-Joon
contents Binarization, which converts weight parameters to binary values, has emerged as an effective strategy to reduce the size of large language models (LLMs). However, typical binarization techniques significantly diminish linguistic effectiveness of LLMs. To address this issue, we introduce a novel binarization technique called Mixture of Scales (BinaryMoS). Unlike conventional methods, BinaryMoS employs multiple scaling experts for binary weights, dynamically merging these experts for each token to adaptively generate scaling factors. This token-adaptive approach boosts the representational power of binarized LLMs by enabling contextual adjustments to the values of binary weights. Moreover, because this adaptive process only involves the scaling factors rather than the entire weight matrix, BinaryMoS maintains compression efficiency similar to traditional static binarization methods. Our experimental results reveal that BinaryMoS surpasses conventional binarization techniques in various natural language processing tasks and even outperforms 2-bit quantization methods, all while maintaining similar model size to static binarization techniques.
format Preprint
id arxiv_https___arxiv_org_abs_2406_12311
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Mixture of Scales: Memory-Efficient Token-Adaptive Binarization for Large Language Models
Jo, Dongwon
Kim, Taesu
Kim, Yulhwa
Kim, Jae-Joon
Machine Learning
Binarization, which converts weight parameters to binary values, has emerged as an effective strategy to reduce the size of large language models (LLMs). However, typical binarization techniques significantly diminish linguistic effectiveness of LLMs. To address this issue, we introduce a novel binarization technique called Mixture of Scales (BinaryMoS). Unlike conventional methods, BinaryMoS employs multiple scaling experts for binary weights, dynamically merging these experts for each token to adaptively generate scaling factors. This token-adaptive approach boosts the representational power of binarized LLMs by enabling contextual adjustments to the values of binary weights. Moreover, because this adaptive process only involves the scaling factors rather than the entire weight matrix, BinaryMoS maintains compression efficiency similar to traditional static binarization methods. Our experimental results reveal that BinaryMoS surpasses conventional binarization techniques in various natural language processing tasks and even outperforms 2-bit quantization methods, all while maintaining similar model size to static binarization techniques.
title Mixture of Scales: Memory-Efficient Token-Adaptive Binarization for Large Language Models
topic Machine Learning
url https://arxiv.org/abs/2406.12311