Improving the Efficiency of a Deep Reinforcement Learning-Based Power Management System for HPC Clusters Using Curriculum Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Budiarjo, Thomas, Pradata, Santana Yuda, Santiyuda, Kadek Gemilang, Amrizal, Muhammad Alfian, Pulungan, Reza, Takizawa, Hiroyuki |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SPARS: A Reinforcement Learning-Enabled Simulator for Power Management in HPC Job Scheduling
by: Amrizal, Muhammad Alfian, et al.
Published: (2025)
by: Amrizal, Muhammad Alfian, et al.
Published: (2025)
A HPC Co-Scheduler with Reinforcement Learning
by: Souza, Abel, et al.
Published: (2024)
by: Souza, Abel, et al.
Published: (2024)
Leveraging Hardware Performance Counters for Predicting Workload Interference in Vector Supercomputers
by: Shubham, et al.
Published: (2024)
by: Shubham, et al.
Published: (2024)
A Reinforcement Learning Based Backfilling Strategy for HPC Batch Jobs
by: Kolker-Hicks, Elliot, et al.
Published: (2024)
by: Kolker-Hicks, Elliot, et al.
Published: (2024)
On the Convergence of Malleability and the HPC PowerStack: Exploiting Dynamism in Over-Provisioned and Power-Constrained HPC Systems
by: Arima, Eishi, et al.
Published: (2024)
by: Arima, Eishi, et al.
Published: (2024)
Solutions for Distributed Memory Access Mechanism on HPC Clusters
by: Meizner, Jan, et al.
Published: (2025)
by: Meizner, Jan, et al.
Published: (2025)
Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters
by: Jain, Rutwik, et al.
Published: (2026)
by: Jain, Rutwik, et al.
Published: (2026)
Computational Performance and Energy Efficiency of ARM based HPC servers
by: Schirmer, Oskar
Published: (2024)
by: Schirmer, Oskar
Published: (2024)
GreenFaaS: Maximizing Energy Efficiency of HPC Workloads with FaaS
by: Kamatar, Alok, et al.
Published: (2024)
by: Kamatar, Alok, et al.
Published: (2024)
Scalable HPC Job Scheduling and Resource Management in SST
by: Abdurahman, Abubeker, et al.
Published: (2025)
by: Abdurahman, Abubeker, et al.
Published: (2025)
Data Version Management and Machine-Actionable Reproducibility for HPC
by: Knüpfer, Andreas, et al.
Published: (2025)
by: Knüpfer, Andreas, et al.
Published: (2025)
Driving Computational Efficiency in Large-Scale Platforms using HPC Technologies
by: Mendez, Alexander Martinez, et al.
Published: (2026)
by: Mendez, Alexander Martinez, et al.
Published: (2026)
MPI-over-CXL: Enhancing Communication Efficiency in Distributed HPC Systems
by: Kwon, Miryeong, et al.
Published: (2025)
by: Kwon, Miryeong, et al.
Published: (2025)
Resource Optimization with MPI Process Malleability for Dynamic Workloads in HPC Clusters
by: Iserte, Sergio, et al.
Published: (2025)
by: Iserte, Sergio, et al.
Published: (2025)
Evaluating Malleable Job Scheduling in HPC Clusters using Real-World Workloads
by: Zojer, Patrick, et al.
Published: (2026)
by: Zojer, Patrick, et al.
Published: (2026)
Power-Aware Scheduling for Multi-Center HPC Electricity Cost Optimization
by: Hossain, Abrar, et al.
Published: (2025)
by: Hossain, Abrar, et al.
Published: (2025)
Profiling and Modeling of Power Characteristics of Leadership-Scale HPC System Workloads
by: Karimi, Ahmad Maroof, et al.
Published: (2024)
by: Karimi, Ahmad Maroof, et al.
Published: (2024)
Understanding Large-Scale HPC System Behavior Through Cluster-Based Visual Analytics
by: Austin, Allison, et al.
Published: (2026)
by: Austin, Allison, et al.
Published: (2026)
KUBEDIRECT: Unleashing the Full Power of the Cluster Manager for Serverless Computing
by: Qi, Sheng, et al.
Published: (2026)
by: Qi, Sheng, et al.
Published: (2026)
Exploring the Frontiers of Energy Efficiency using Power Management at System Scale
by: Karimi, Ahmad Maroof, et al.
Published: (2024)
by: Karimi, Ahmad Maroof, et al.
Published: (2024)
Attack Graph Generation on HPC Clusters
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
Integrating and Characterizing HPC Task Runtime Systems for hybrid AI-HPC workloads
by: Merzky, Andre, et al.
Published: (2025)
by: Merzky, Andre, et al.
Published: (2025)
Scale: Deep Reinforcement Learning for Container Scheduling in Serverless Edge Computing
by: Chen, Chen, et al.
Published: (2026)
by: Chen, Chen, et al.
Published: (2026)
Efficient Hierarchical Storage Management Framework Empowered by Reinforcement Learning
by: Zhang, Tianru, et al.
Published: (2022)
by: Zhang, Tianru, et al.
Published: (2022)
DiT-HC: Enabling Efficient Training of Visual Generation Model DiT on HPC-oriented CPU Cluster
by: Zhang, Jinxiao, et al.
Published: (2026)
by: Zhang, Jinxiao, et al.
Published: (2026)
HPC with Enhanced User Separation
by: Prout, Andrew, et al.
Published: (2024)
by: Prout, Andrew, et al.
Published: (2024)
Analysis of the carbon footprint of HPC
by: Benhari, Abdessalam, et al.
Published: (2025)
by: Benhari, Abdessalam, et al.
Published: (2025)
DRPC: Distributed Reinforcement Learning Approach for Scalable Resource Provisioning in Container-based Clusters
by: Bai, Haoyu, et al.
Published: (2024)
by: Bai, Haoyu, et al.
Published: (2024)
MRSch: Multi-Resource Scheduling for HPC
by: Li, Boyang, et al.
Published: (2024)
by: Li, Boyang, et al.
Published: (2024)
Modernizing an Operational Real-time Tsunami Simulator to Support Diverse Hardware Platforms
by: Takahashi, Keichi, et al.
Published: (2024)
by: Takahashi, Keichi, et al.
Published: (2024)
NL-CPS: Reinforcement Learning-Based Kubernetes Control Plane Placement in Multi-Region Clusters
by: Alam, Sajid, et al.
Published: (2026)
by: Alam, Sajid, et al.
Published: (2026)
Resilient Packet Forwarding: A Reinforcement Learning Approach to Routing in Gaussian Interconnected Networks with Clustered Faults
by: Charrwi, Mohammad Walid, et al.
Published: (2025)
by: Charrwi, Mohammad Walid, et al.
Published: (2025)
HiRL: Hierarchical Reinforcement Learning for Coordinated Resource Management in Heterogeneous Edge Computing
by: Zhu, Jianyong, et al.
Published: (2026)
by: Zhu, Jianyong, et al.
Published: (2026)
UNR: Unified Notifiable RMA Library for HPC
by: Feng, Guangnan, et al.
Published: (2024)
by: Feng, Guangnan, et al.
Published: (2024)
An Elastic Job Scheduler for HPC Applications on the Cloud
by: Bhosale, Aditya, et al.
Published: (2025)
by: Bhosale, Aditya, et al.
Published: (2025)
Sarus Suite: Cloud-native Containers for HPC
by: Madonna, Alberto, et al.
Published: (2026)
by: Madonna, Alberto, et al.
Published: (2026)
Wilkins: HPC In Situ Workflows Made Easy
by: Yildiz, Orcun, et al.
Published: (2024)
by: Yildiz, Orcun, et al.
Published: (2024)
Characterizing the Impact of Congestion in Modern HPC Interconnects
by: Piarulli, Lorenzo, et al.
Published: (2026)
by: Piarulli, Lorenzo, et al.
Published: (2026)
Towards Energy Efficient Co-Scheduling in HPC
by: Zheng, Zhong, et al.
Published: (2026)
by: Zheng, Zhong, et al.
Published: (2026)
Leveraging Teaching on Demand: Approaching HPC to Undergrads
by: Catalán, S., et al.
Published: (2026)
by: Catalán, S., et al.
Published: (2026)
Similar Items
-
SPARS: A Reinforcement Learning-Enabled Simulator for Power Management in HPC Job Scheduling
by: Amrizal, Muhammad Alfian, et al.
Published: (2025) -
A HPC Co-Scheduler with Reinforcement Learning
by: Souza, Abel, et al.
Published: (2024) -
Leveraging Hardware Performance Counters for Predicting Workload Interference in Vector Supercomputers
by: Shubham, et al.
Published: (2024) -
A Reinforcement Learning Based Backfilling Strategy for HPC Batch Jobs
by: Kolker-Hicks, Elliot, et al.
Published: (2024) -
On the Convergence of Malleability and the HPC PowerStack: Exploiting Dynamism in Over-Provisioned and Power-Constrained HPC Systems
by: Arima, Eishi, et al.
Published: (2024)