Saved in:
Bibliographic Details
Main Authors: Canakci, Burcu, Liu, Junyi, Wu, Xingbo, Cheriere, Nathanaël, Costa, Paolo, Legtchenko, Sergey, Narayanan, Dushyanth, Rowstron, Ant
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2501.10187
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913811602079744
author Canakci, Burcu
Liu, Junyi
Wu, Xingbo
Cheriere, Nathanaël
Costa, Paolo
Legtchenko, Sergey
Narayanan, Dushyanth
Rowstron, Ant
author_facet Canakci, Burcu
Liu, Junyi
Wu, Xingbo
Cheriere, Nathanaël
Costa, Paolo
Legtchenko, Sergey
Narayanan, Dushyanth
Rowstron, Ant
contents To match the blooming demand of generative AI workloads, GPU designers have so far been trying to pack more and more compute and memory into single complex and expensive packages. However, there is growing uncertainty about the scalability of individual GPUs and thus AI clusters, as state-of-the-art GPUs are already displaying packaging, yield, and cooling limitations. We propose to rethink the design and scaling of AI clusters through efficiently-connected large clusters of Lite-GPUs, GPUs with single, small dies and a fraction of the capabilities of larger GPUs. We think recent advances in co-packaged optics can enable distributing AI workloads onto many Lite-GPUs through high bandwidth and efficient communication. In this paper, we present the key benefits of Lite-GPUs on manufacturing cost, blast radius, yield, and power efficiency; and discuss systems opportunities and challenges around resource, workload, memory, and network management.
format Preprint
id arxiv_https___arxiv_org_abs_2501_10187
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Good things come in small packages: Should we build AI clusters with Lite-GPUs?
Canakci, Burcu
Liu, Junyi
Wu, Xingbo
Cheriere, Nathanaël
Costa, Paolo
Legtchenko, Sergey
Narayanan, Dushyanth
Rowstron, Ant
Hardware Architecture
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
To match the blooming demand of generative AI workloads, GPU designers have so far been trying to pack more and more compute and memory into single complex and expensive packages. However, there is growing uncertainty about the scalability of individual GPUs and thus AI clusters, as state-of-the-art GPUs are already displaying packaging, yield, and cooling limitations. We propose to rethink the design and scaling of AI clusters through efficiently-connected large clusters of Lite-GPUs, GPUs with single, small dies and a fraction of the capabilities of larger GPUs. We think recent advances in co-packaged optics can enable distributing AI workloads onto many Lite-GPUs through high bandwidth and efficient communication. In this paper, we present the key benefits of Lite-GPUs on manufacturing cost, blast radius, yield, and power efficiency; and discuss systems opportunities and challenges around resource, workload, memory, and network management.
title Good things come in small packages: Should we build AI clusters with Lite-GPUs?
topic Hardware Architecture
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2501.10187