Scaling Intelligence: Designing Data Centers for Next-Gen Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tithi, Jesmin Jahan, Wu, Hanjiang, Abuhatzera, Avishaii, Petrini, Fabrizio
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912573794811904
author Tithi, Jesmin Jahan
Wu, Hanjiang
Abuhatzera, Avishaii
Petrini, Fabrizio
author_facet Tithi, Jesmin Jahan
Wu, Hanjiang
Abuhatzera, Avishaii
Petrini, Fabrizio
contents The explosive growth of Large Language Models (LLMs), such as GPT-4 with 1.8 trillion parameters, demands a fundamental rethinking of data center architecture to ensure scalability, efficiency, and cost-effectiveness. Our work provides a comprehensive co-design framework that jointly explores FLOPS, HBM bandwidth and capacity, multiple network topologies (two-tier vs. FullFlat optical), the size of the scale-out domain, and popular parallelism/optimization strategies used in LLMs. We introduce and evaluate FullFlat network architectures, which provide uniform high-bandwidth, low-latency connectivity between all nodes, and demonstrate their transformative impact on performance and scalability. Through detailed sensitivity analyses, we quantify the benefits of overlapping compute and communication, leveraging hardware-accelerated collectives, widening the scale-out domain, and increasing memory capacity. Our study spans both sparse (mixture of experts) and dense transformer-based LLMs, revealing how system design choices affect Model FLOPS Utilization (MFU = Model FLOPS per token * Observed tokens per second / Peak FLOPS of the hardware) and overall throughput. For the co-design study, we utilized an analytical performance modeling tool capable of predicting LLM runtime within 10% of real-world measurements. Our findings offer actionable insights and a practical roadmap for designing AI data centers that can efficiently support trillion-parameter models, reduce optimization complexity, and sustain the rapid evolution of AI capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2506_15006
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scaling Intelligence: Designing Data Centers for Next-Gen Language Models
Tithi, Jesmin Jahan
Wu, Hanjiang
Abuhatzera, Avishaii
Petrini, Fabrizio
Hardware Architecture
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Emerging Technologies
Performance
The explosive growth of Large Language Models (LLMs), such as GPT-4 with 1.8 trillion parameters, demands a fundamental rethinking of data center architecture to ensure scalability, efficiency, and cost-effectiveness. Our work provides a comprehensive co-design framework that jointly explores FLOPS, HBM bandwidth and capacity, multiple network topologies (two-tier vs. FullFlat optical), the size of the scale-out domain, and popular parallelism/optimization strategies used in LLMs. We introduce and evaluate FullFlat network architectures, which provide uniform high-bandwidth, low-latency connectivity between all nodes, and demonstrate their transformative impact on performance and scalability. Through detailed sensitivity analyses, we quantify the benefits of overlapping compute and communication, leveraging hardware-accelerated collectives, widening the scale-out domain, and increasing memory capacity. Our study spans both sparse (mixture of experts) and dense transformer-based LLMs, revealing how system design choices affect Model FLOPS Utilization (MFU = Model FLOPS per token * Observed tokens per second / Peak FLOPS of the hardware) and overall throughput. For the co-design study, we utilized an analytical performance modeling tool capable of predicting LLM runtime within 10% of real-world measurements. Our findings offer actionable insights and a practical roadmap for designing AI data centers that can efficiently support trillion-parameter models, reduce optimization complexity, and sustain the rapid evolution of AI capabilities.
title Scaling Intelligence: Designing Data Centers for Next-Gen Language Models
topic Hardware Architecture
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Emerging Technologies
Performance
url https://arxiv.org/abs/2506.15006