ATLAHS: An Application-centric Network Simulator Toolchain for AI, HPC, and Distributed Storage

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Siyuan, Bonato, Tommaso, Hu, Zhiyi, Jordan, Pasquale, Chen, Tiancheng, Hoefler, Torsten
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912374396551168
author Shen, Siyuan
Bonato, Tommaso
Hu, Zhiyi
Jordan, Pasquale
Chen, Tiancheng
Hoefler, Torsten
author_facet Shen, Siyuan
Bonato, Tommaso
Hu, Zhiyi
Jordan, Pasquale
Chen, Tiancheng
Hoefler, Torsten
contents Network simulators play a crucial role in evaluating the performance of large-scale systems. However, existing simulators rely heavily on synthetic microbenchmarks or narrowly focus on specific domains, limiting their ability to provide comprehensive performance insights. In this work, we introduce ATLAHS, a flexible, extensible, and open-source toolchain designed to trace real-world applications and accurately simulate their workloads. ATLAHS leverages the GOAL format to model communication and computation patterns in AI, HPC, and distributed storage applications. It supports multiple network simulation backends and handles multi-job and multi-tenant scenarios. Through extensive validation, we demonstrate that ATLAHS achieves high accuracy in simulating realistic workloads (consistently less than 5% error), while significantly outperforming AstraSim, the current state-of-the-art AI systems simulator, in terms of simulation runtime and trace size efficiency. We further illustrate ATLAHS's utility via detailed case studies, highlighting the impact of congestion control algorithms on the performance of distributed storage systems, as well as the influence of job-placement strategies on application runtimes.
format Preprint
id arxiv_https___arxiv_org_abs_2505_08936
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ATLAHS: An Application-centric Network Simulator Toolchain for AI, HPC, and Distributed Storage
Shen, Siyuan
Bonato, Tommaso
Hu, Zhiyi
Jordan, Pasquale
Chen, Tiancheng
Hoefler, Torsten
Distributed, Parallel, and Cluster Computing
I.6.3
Network simulators play a crucial role in evaluating the performance of large-scale systems. However, existing simulators rely heavily on synthetic microbenchmarks or narrowly focus on specific domains, limiting their ability to provide comprehensive performance insights. In this work, we introduce ATLAHS, a flexible, extensible, and open-source toolchain designed to trace real-world applications and accurately simulate their workloads. ATLAHS leverages the GOAL format to model communication and computation patterns in AI, HPC, and distributed storage applications. It supports multiple network simulation backends and handles multi-job and multi-tenant scenarios. Through extensive validation, we demonstrate that ATLAHS achieves high accuracy in simulating realistic workloads (consistently less than 5% error), while significantly outperforming AstraSim, the current state-of-the-art AI systems simulator, in terms of simulation runtime and trace size efficiency. We further illustrate ATLAHS's utility via detailed case studies, highlighting the impact of congestion control algorithms on the performance of distributed storage systems, as well as the influence of job-placement strategies on application runtimes.
title ATLAHS: An Application-centric Network Simulator Toolchain for AI, HPC, and Distributed Storage
topic Distributed, Parallel, and Cluster Computing
I.6.3
url https://arxiv.org/abs/2505.08936