The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ma, Shuming, Wang, Hongyu, Ma, Lingxiao, Wang, Lei, Wang, Wenhui, Huang, Shaohan, Dong, Li, Wang, Ruiping, Xue, Jilong, Wei, Furu
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917599632162816
author Ma, Shuming
Wang, Hongyu
Ma, Lingxiao
Wang, Lei
Wang, Wenhui
Huang, Shaohan
Dong, Li
Wang, Ruiping
Xue, Jilong
Wei, Furu
author_facet Ma, Shuming
Wang, Hongyu
Ma, Lingxiao
Wang, Lei
Wang, Wenhui
Huang, Shaohan
Dong, Li
Wang, Ruiping
Xue, Jilong
Wei, Furu
contents Recent research, such as BitNet, is paving the way for a new era of 1-bit Large Language Models (LLMs). In this work, we introduce a 1-bit LLM variant, namely BitNet b1.58, in which every single parameter (or weight) of the LLM is ternary {-1, 0, 1}. It matches the full-precision (i.e., FP16 or BF16) Transformer LLM with the same model size and training tokens in terms of both perplexity and end-task performance, while being significantly more cost-effective in terms of latency, memory, throughput, and energy consumption. More profoundly, the 1.58-bit LLM defines a new scaling law and recipe for training new generations of LLMs that are both high-performance and cost-effective. Furthermore, it enables a new computation paradigm and opens the door for designing specific hardware optimized for 1-bit LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2402_17764
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
Ma, Shuming
Wang, Hongyu
Ma, Lingxiao
Wang, Lei
Wang, Wenhui
Huang, Shaohan
Dong, Li
Wang, Ruiping
Xue, Jilong
Wei, Furu
Computation and Language
Machine Learning
Recent research, such as BitNet, is paving the way for a new era of 1-bit Large Language Models (LLMs). In this work, we introduce a 1-bit LLM variant, namely BitNet b1.58, in which every single parameter (or weight) of the LLM is ternary {-1, 0, 1}. It matches the full-precision (i.e., FP16 or BF16) Transformer LLM with the same model size and training tokens in terms of both perplexity and end-task performance, while being significantly more cost-effective in terms of latency, memory, throughput, and energy consumption. More profoundly, the 1.58-bit LLM defines a new scaling law and recipe for training new generations of LLMs that are both high-performance and cost-effective. Furthermore, it enables a new computation paradigm and opens the door for designing specific hardware optimized for 1-bit LLMs.
title The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2402.17764