SPARS: A Reinforcement Learning-Enabled Simulator for Power Management in HPC Job Scheduling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Amrizal, Muhammad Alfian, Prasasta, Raka Satya, Pradata, Santana Yuda, Santiyuda, Kadek Gemilang, Pulungan, Reza, Takizawa, Hiroyuki
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916046128021504
author Amrizal, Muhammad Alfian
Prasasta, Raka Satya
Pradata, Santana Yuda
Santiyuda, Kadek Gemilang
Pulungan, Reza
Takizawa, Hiroyuki
author_facet Amrizal, Muhammad Alfian
Prasasta, Raka Satya
Pradata, Santana Yuda
Santiyuda, Kadek Gemilang
Pulungan, Reza
Takizawa, Hiroyuki
contents High-performance computing (HPC) systems consume enormous amounts of energy, with idle nodes as a major source of energy waste. Powering down idle nodes can mitigate this problem, but long boot/shutdown delays can introduce significant queueing penalties if transitions are poorly timed. To address this trade-off, we present SPARS, a reinforcement learning-enabled simulator for power management in HPC job scheduling. SPARS integrates job scheduling and node power-state management within a discrete-event simulation framework. It supports traditional scheduling policies such as First Come First Serve and EASY Backfilling, along with enhanced variants that employ reinforcement learning agents to dynamically decide when nodes should be powered on or off. Users can configure workloads and platforms in JSON format, specifying job arrivals, execution times, node power models, and transition delays. The simulator records comprehensive metrics-including energy usage, wasted power, job waiting times, and node utilization-and provides Gantt chart visualizations to analyze scheduling dynamics and power transitions. Unlike widely used Batsim-based frameworks that rely on heavy inter-process communication, SPARS provides lightweight event handling and consistent simulation results, making experiments easier to reproduce and extend. Its modular design allows new scheduling heuristics or learning algorithms to be integrated with minimal effort. By providing a flexible, reproducible, and extensible platform, SPARS enables researchers and practitioners to systematically evaluate power-aware scheduling strategies, explore the trade-offs between energy efficiency and performance, and accelerate the development of sustainable HPC operations.
format Preprint
id arxiv_https___arxiv_org_abs_2512_13268
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SPARS: A Reinforcement Learning-Enabled Simulator for Power Management in HPC Job Scheduling
Amrizal, Muhammad Alfian
Prasasta, Raka Satya
Pradata, Santana Yuda
Santiyuda, Kadek Gemilang
Pulungan, Reza
Takizawa, Hiroyuki
Distributed, Parallel, and Cluster Computing
High-performance computing (HPC) systems consume enormous amounts of energy, with idle nodes as a major source of energy waste. Powering down idle nodes can mitigate this problem, but long boot/shutdown delays can introduce significant queueing penalties if transitions are poorly timed. To address this trade-off, we present SPARS, a reinforcement learning-enabled simulator for power management in HPC job scheduling. SPARS integrates job scheduling and node power-state management within a discrete-event simulation framework. It supports traditional scheduling policies such as First Come First Serve and EASY Backfilling, along with enhanced variants that employ reinforcement learning agents to dynamically decide when nodes should be powered on or off. Users can configure workloads and platforms in JSON format, specifying job arrivals, execution times, node power models, and transition delays. The simulator records comprehensive metrics-including energy usage, wasted power, job waiting times, and node utilization-and provides Gantt chart visualizations to analyze scheduling dynamics and power transitions. Unlike widely used Batsim-based frameworks that rely on heavy inter-process communication, SPARS provides lightweight event handling and consistent simulation results, making experiments easier to reproduce and extend. Its modular design allows new scheduling heuristics or learning algorithms to be integrated with minimal effort. By providing a flexible, reproducible, and extensible platform, SPARS enables researchers and practitioners to systematically evaluate power-aware scheduling strategies, explore the trade-offs between energy efficiency and performance, and accelerate the development of sustainable HPC operations.
title SPARS: A Reinforcement Learning-Enabled Simulator for Power Management in HPC Job Scheduling
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2512.13268