When are 1.58 bits enough? A Bottom-up Exploration of BitNet Quantization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nielsen, Jacob, Galke, Lukas, Schneider-Kamp, Peter
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912112933076992
author Nielsen, Jacob
Galke, Lukas
Schneider-Kamp, Peter
author_facet Nielsen, Jacob
Galke, Lukas
Schneider-Kamp, Peter
contents Contemporary machine learning models, such as language models, are powerful, but come with immense resource requirements both at training and inference time. It has been shown that decoder-only language models can be trained to a competitive state with ternary weights (1.58 bits per weight), facilitating efficient inference. Here, we start our exploration with non-transformer model architectures, investigating 1.58-bit training for multi-layer perceptrons and graph neural networks. Then, we explore 1.58-bit training in other transformer-based language models, namely encoder-only and encoder-decoder models. Our results show that in all of these settings, 1.58-bit training is on par with or sometimes even better than the standard 32/16-bit models.
format Preprint
id arxiv_https___arxiv_org_abs_2411_05882
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle When are 1.58 bits enough? A Bottom-up Exploration of BitNet Quantization
Nielsen, Jacob
Galke, Lukas
Schneider-Kamp, Peter
Machine Learning
Computation and Language
Contemporary machine learning models, such as language models, are powerful, but come with immense resource requirements both at training and inference time. It has been shown that decoder-only language models can be trained to a competitive state with ternary weights (1.58 bits per weight), facilitating efficient inference. Here, we start our exploration with non-transformer model architectures, investigating 1.58-bit training for multi-layer perceptrons and graph neural networks. Then, we explore 1.58-bit training in other transformer-based language models, namely encoder-only and encoder-decoder models. Our results show that in all of these settings, 1.58-bit training is on par with or sometimes even better than the standard 32/16-bit models.
title When are 1.58 bits enough? A Bottom-up Exploration of BitNet Quantization
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2411.05882