AI and Memory Wall
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916170082287616 |
|---|---|
| author | Gholami, Amir Yao, Zhewei Kim, Sehoon Hooper, Coleman Mahoney, Michael W. Keutzer, Kurt |
| author_facet | Gholami, Amir Yao, Zhewei Kim, Sehoon Hooper, Coleman Mahoney, Michael W. Keutzer, Kurt |
| contents | The availability of unprecedented unsupervised training data, along with neural scaling laws, has resulted in an unprecedented surge in model size and compute requirements for serving/training LLMs. However, the main performance bottleneck is increasingly shifting to memory bandwidth. Over the past 20 years, peak server hardware FLOPS has been scaling at 3.0x/2yrs, outpacing the growth of DRAM and interconnect bandwidth, which have only scaled at 1.6 and 1.4 times every 2 years, respectively. This disparity has made memory, rather than compute, the primary bottleneck in AI applications, particularly in serving. Here, we analyze encoder and decoder Transformer models and show how memory bandwidth can become the dominant bottleneck for decoder models. We argue for a redesign in model architecture, training, and deployment strategies to overcome this memory limitation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2403_14123 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | AI and Memory Wall Gholami, Amir Yao, Zhewei Kim, Sehoon Hooper, Coleman Mahoney, Michael W. Keutzer, Kurt Machine Learning Hardware Architecture Distributed, Parallel, and Cluster Computing The availability of unprecedented unsupervised training data, along with neural scaling laws, has resulted in an unprecedented surge in model size and compute requirements for serving/training LLMs. However, the main performance bottleneck is increasingly shifting to memory bandwidth. Over the past 20 years, peak server hardware FLOPS has been scaling at 3.0x/2yrs, outpacing the growth of DRAM and interconnect bandwidth, which have only scaled at 1.6 and 1.4 times every 2 years, respectively. This disparity has made memory, rather than compute, the primary bottleneck in AI applications, particularly in serving. Here, we analyze encoder and decoder Transformer models and show how memory bandwidth can become the dominant bottleneck for decoder models. We argue for a redesign in model architecture, training, and deployment strategies to overcome this memory limitation. |
| title | AI and Memory Wall |
| topic | Machine Learning Hardware Architecture Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2403.14123 |