Scalable MatMul-free Language Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Rui-Jie, Zhang, Yu, Abreu, Steven, Sifferman, Ethan, Sheaves, Tyler, Wang, Yiqiao, Richmond, Dustin, Shrestha, Sumit Bam, Zhou, Peng, Eshraghian, Jason K.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911077778849792
author Zhu, Rui-Jie
Zhang, Yu
Abreu, Steven
Sifferman, Ethan
Sheaves, Tyler
Wang, Yiqiao
Richmond, Dustin
Shrestha, Sumit Bam
Zhou, Peng
Eshraghian, Jason K.
author_facet Zhu, Rui-Jie
Zhang, Yu
Abreu, Steven
Sifferman, Ethan
Sheaves, Tyler
Wang, Yiqiao
Richmond, Dustin
Shrestha, Sumit Bam
Zhou, Peng
Eshraghian, Jason K.
contents Large Language Models (LLMs) have fundamentally altered how we approach scaling in machine learning. However, these models pose substantial computational and memory challenges, primarily due to the reliance on matrix multiplication (MatMul) within their attention and feed-forward (FFN) layers. We demonstrate that MatMul operations can be eliminated from LLMs while maintaining strong performance, even at billion-parameter scales. Our MatMul-free models, tested on models up to 2.7B parameters, are comparable to state-of-the-art pre-trained Transformers, and the performance gap narrows as model size increases. Our approach yields significant memory savings: a GPU-efficient implementation reduces memory consumption by up to 61% during training and over 10x during inference. When adapted for a multi-chip neuromorphic system, the model leverages asynchronous processing to achieve 4x higher throughput with 10x less energy than edge GPUs.
format Preprint
id arxiv_https___arxiv_org_abs_2406_02528
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Scalable MatMul-free Language Modeling
Zhu, Rui-Jie
Zhang, Yu
Abreu, Steven
Sifferman, Ethan
Sheaves, Tyler
Wang, Yiqiao
Richmond, Dustin
Shrestha, Sumit Bam
Zhou, Peng
Eshraghian, Jason K.
Computation and Language
Large Language Models (LLMs) have fundamentally altered how we approach scaling in machine learning. However, these models pose substantial computational and memory challenges, primarily due to the reliance on matrix multiplication (MatMul) within their attention and feed-forward (FFN) layers. We demonstrate that MatMul operations can be eliminated from LLMs while maintaining strong performance, even at billion-parameter scales. Our MatMul-free models, tested on models up to 2.7B parameters, are comparable to state-of-the-art pre-trained Transformers, and the performance gap narrows as model size increases. Our approach yields significant memory savings: a GPU-efficient implementation reduces memory consumption by up to 61% during training and over 10x during inference. When adapted for a multi-chip neuromorphic system, the model leverages asynchronous processing to achieve 4x higher throughput with 10x less energy than edge GPUs.
title Scalable MatMul-free Language Modeling
topic Computation and Language
url https://arxiv.org/abs/2406.02528