The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866917599632162816 |
|---|---|
| author | Ma, Shuming Wang, Hongyu Ma, Lingxiao Wang, Lei Wang, Wenhui Huang, Shaohan Dong, Li Wang, Ruiping Xue, Jilong Wei, Furu |
| author_facet | Ma, Shuming Wang, Hongyu Ma, Lingxiao Wang, Lei Wang, Wenhui Huang, Shaohan Dong, Li Wang, Ruiping Xue, Jilong Wei, Furu |
| contents | Recent research, such as BitNet, is paving the way for a new era of 1-bit Large Language Models (LLMs). In this work, we introduce a 1-bit LLM variant, namely BitNet b1.58, in which every single parameter (or weight) of the LLM is ternary {-1, 0, 1}. It matches the full-precision (i.e., FP16 or BF16) Transformer LLM with the same model size and training tokens in terms of both perplexity and end-task performance, while being significantly more cost-effective in terms of latency, memory, throughput, and energy consumption. More profoundly, the 1.58-bit LLM defines a new scaling law and recipe for training new generations of LLMs that are both high-performance and cost-effective. Furthermore, it enables a new computation paradigm and opens the door for designing specific hardware optimized for 1-bit LLMs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2402_17764 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits Ma, Shuming Wang, Hongyu Ma, Lingxiao Wang, Lei Wang, Wenhui Huang, Shaohan Dong, Li Wang, Ruiping Xue, Jilong Wei, Furu Computation and Language Machine Learning Recent research, such as BitNet, is paving the way for a new era of 1-bit Large Language Models (LLMs). In this work, we introduce a 1-bit LLM variant, namely BitNet b1.58, in which every single parameter (or weight) of the LLM is ternary {-1, 0, 1}. It matches the full-precision (i.e., FP16 or BF16) Transformer LLM with the same model size and training tokens in terms of both perplexity and end-task performance, while being significantly more cost-effective in terms of latency, memory, throughput, and energy consumption. More profoundly, the 1.58-bit LLM defines a new scaling law and recipe for training new generations of LLMs that are both high-performance and cost-effective. Furthermore, it enables a new computation paradigm and opens the door for designing specific hardware optimized for 1-bit LLMs. |
| title | The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2402.17764 |