Trace Replay Simulation of MIT SuperCloud for Studying Optimal Sustainability Policies

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Brewer, Wesley, Maiterth, Matthias, Fay, Damien
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909798471041024
author Brewer, Wesley
Maiterth, Matthias
Fay, Damien
author_facet Brewer, Wesley
Maiterth, Matthias
Fay, Damien
contents The rapid growth of AI supercomputing is creating unprecedented power demands, with next-generation GPU datacenters requiring hundreds of megawatts and producing fast, large swings in consumption. To address the resulting challenges for utilities and system operators, we extend ExaDigiT, an open-source digital twin framework for modeling power, cooling, and scheduling of supercomputers. Originally developed for replaying traces from leadership-class HPC systems, ExaDigiT now incorporates heterogeneity, multi-tenancy, and cloud-scale workloads. In this work, we focus on trace replay and rescheduling of jobs on the MIT SuperCloud TX-GAIA system to enable reinforcement learning (RL)-based experimentation with sustainability policies. The RAPS module provides a simulation environment with detailed power and performance statistics, supporting the study of scheduling strategies, incentive structures, and hardware/software prototyping. Preliminary RL experiments using Proximal Policy Optimization demonstrate the feasibility of learning energy-aware scheduling decisions, highlighting ExaDigiT's potential as a platform for exploring optimal policies to improve throughput, efficiency, and sustainability.
format Preprint
id arxiv_https___arxiv_org_abs_2509_16513
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Trace Replay Simulation of MIT SuperCloud for Studying Optimal Sustainability Policies
Brewer, Wesley
Maiterth, Matthias
Fay, Damien
Distributed, Parallel, and Cluster Computing
The rapid growth of AI supercomputing is creating unprecedented power demands, with next-generation GPU datacenters requiring hundreds of megawatts and producing fast, large swings in consumption. To address the resulting challenges for utilities and system operators, we extend ExaDigiT, an open-source digital twin framework for modeling power, cooling, and scheduling of supercomputers. Originally developed for replaying traces from leadership-class HPC systems, ExaDigiT now incorporates heterogeneity, multi-tenancy, and cloud-scale workloads. In this work, we focus on trace replay and rescheduling of jobs on the MIT SuperCloud TX-GAIA system to enable reinforcement learning (RL)-based experimentation with sustainability policies. The RAPS module provides a simulation environment with detailed power and performance statistics, supporting the study of scheduling strategies, incentive structures, and hardware/software prototyping. Preliminary RL experiments using Proximal Policy Optimization demonstrate the feasibility of learning energy-aware scheduling decisions, highlighting ExaDigiT's potential as a platform for exploring optimal policies to improve throughput, efficiency, and sustainability.
title Trace Replay Simulation of MIT SuperCloud for Studying Optimal Sustainability Policies
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2509.16513