Matmul or No Matmul in the Era of 1-bit LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Malekar, Jinendra, Elbtity, Mohammed E., Zand, Ramtin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909300195065856
author Malekar, Jinendra
Elbtity, Mohammed E.
Zand, Ramtin
author_facet Malekar, Jinendra
Elbtity, Mohammed E.
Zand, Ramtin
contents The advent of 1-bit large language models (LLMs) has attracted considerable attention and opened up new research opportunities. However, 1-bit LLMs only improve a fraction of models by applying extreme quantization to the projection layers while leaving attention heads unchanged. Therefore, to avoid fundamentally wrong choices of goals in future research, it is crucial to understand the actual improvements in computation and memory usage that 1-bit LLMs can deliver. In this work, we present an adaptation of Amdahl's Law tailored for the 1-bit LLM context, which illustrates how partial improvements in 1-bit LLMs impact overall model performance. Through extensive experiments, we uncover key nuances across different model architectures and hardware configurations, offering a roadmap for future research in the era of 1-bit LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2408_11939
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Matmul or No Matmul in the Era of 1-bit LLMs
Malekar, Jinendra
Elbtity, Mohammed E.
Zand, Ramtin
Artificial Intelligence
Machine Learning
The advent of 1-bit large language models (LLMs) has attracted considerable attention and opened up new research opportunities. However, 1-bit LLMs only improve a fraction of models by applying extreme quantization to the projection layers while leaving attention heads unchanged. Therefore, to avoid fundamentally wrong choices of goals in future research, it is crucial to understand the actual improvements in computation and memory usage that 1-bit LLMs can deliver. In this work, we present an adaptation of Amdahl's Law tailored for the 1-bit LLM context, which illustrates how partial improvements in 1-bit LLMs impact overall model performance. Through extensive experiments, we uncover key nuances across different model architectures and hardware configurations, offering a roadmap for future research in the era of 1-bit LLMs.
title Matmul or No Matmul in the Era of 1-bit LLMs
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2408.11939