AI and Memory Wall

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gholami, Amir, Yao, Zhewei, Kim, Sehoon, Hooper, Coleman, Mahoney, Michael W., Keutzer, Kurt
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916170082287616
author Gholami, Amir
Yao, Zhewei
Kim, Sehoon
Hooper, Coleman
Mahoney, Michael W.
Keutzer, Kurt
author_facet Gholami, Amir
Yao, Zhewei
Kim, Sehoon
Hooper, Coleman
Mahoney, Michael W.
Keutzer, Kurt
contents The availability of unprecedented unsupervised training data, along with neural scaling laws, has resulted in an unprecedented surge in model size and compute requirements for serving/training LLMs. However, the main performance bottleneck is increasingly shifting to memory bandwidth. Over the past 20 years, peak server hardware FLOPS has been scaling at 3.0x/2yrs, outpacing the growth of DRAM and interconnect bandwidth, which have only scaled at 1.6 and 1.4 times every 2 years, respectively. This disparity has made memory, rather than compute, the primary bottleneck in AI applications, particularly in serving. Here, we analyze encoder and decoder Transformer models and show how memory bandwidth can become the dominant bottleneck for decoder models. We argue for a redesign in model architecture, training, and deployment strategies to overcome this memory limitation.
format Preprint
id arxiv_https___arxiv_org_abs_2403_14123
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AI and Memory Wall
Gholami, Amir
Yao, Zhewei
Kim, Sehoon
Hooper, Coleman
Mahoney, Michael W.
Keutzer, Kurt
Machine Learning
Hardware Architecture
Distributed, Parallel, and Cluster Computing
The availability of unprecedented unsupervised training data, along with neural scaling laws, has resulted in an unprecedented surge in model size and compute requirements for serving/training LLMs. However, the main performance bottleneck is increasingly shifting to memory bandwidth. Over the past 20 years, peak server hardware FLOPS has been scaling at 3.0x/2yrs, outpacing the growth of DRAM and interconnect bandwidth, which have only scaled at 1.6 and 1.4 times every 2 years, respectively. This disparity has made memory, rather than compute, the primary bottleneck in AI applications, particularly in serving. Here, we analyze encoder and decoder Transformer models and show how memory bandwidth can become the dominant bottleneck for decoder models. We argue for a redesign in model architecture, training, and deployment strategies to overcome this memory limitation.
title AI and Memory Wall
topic Machine Learning
Hardware Architecture
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2403.14123