On the Convergence of Malleability and the HPC PowerStack: Exploiting Dynamism in Over-Provisioned and Power-Constrained HPC Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Arima, Eishi, Comprés, Isaías A., Schulz, Martin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Orchestrated Co-scheduling, Resource Partitioning, and Power Capping on CPU-GPU Heterogeneous Systems via Machine Learning
by: Saba, Issa, et al.
Published: (2024)
by: Saba, Issa, et al.
Published: (2024)
Resource Optimization with MPI Process Malleability for Dynamic Workloads in HPC Clusters
by: Iserte, Sergio, et al.
Published: (2025)
by: Iserte, Sergio, et al.
Published: (2025)
MPI Malleability Validation under Replayed Real-World HPC Conditions
by: Iserte, S., et al.
Published: (2026)
by: Iserte, S., et al.
Published: (2026)
Optimizing Hardware Resource Partitioning and Job Allocations on Modern GPUs under Power Caps
by: Arima, Eishi, et al.
Published: (2024)
by: Arima, Eishi, et al.
Published: (2024)
Evaluating Malleable Job Scheduling in HPC Clusters using Real-World Workloads
by: Zojer, Patrick, et al.
Published: (2026)
by: Zojer, Patrick, et al.
Published: (2026)
Profiling and Modeling of Power Characteristics of Leadership-Scale HPC System Workloads
by: Karimi, Ahmad Maroof, et al.
Published: (2024)
by: Karimi, Ahmad Maroof, et al.
Published: (2024)
Power-Aware Scheduling for Multi-Center HPC Electricity Cost Optimization
by: Hossain, Abrar, et al.
Published: (2025)
by: Hossain, Abrar, et al.
Published: (2025)
Integrating and Characterizing HPC Task Runtime Systems for hybrid AI-HPC workloads
by: Merzky, Andre, et al.
Published: (2025)
by: Merzky, Andre, et al.
Published: (2025)
Introducing JIRIAF: A Virtual Kubelet Integration for Optimizing HPC Resource Provisioning
by: Gyurjyan, Vardan, et al.
Published: (2025)
by: Gyurjyan, Vardan, et al.
Published: (2025)
A Contention-Free Model for Converged Kubernetes on HPC
by: Sochat, Vanessa, et al.
Published: (2024)
by: Sochat, Vanessa, et al.
Published: (2024)
SPARS: A Reinforcement Learning-Enabled Simulator for Power Management in HPC Job Scheduling
by: Amrizal, Muhammad Alfian, et al.
Published: (2025)
by: Amrizal, Muhammad Alfian, et al.
Published: (2025)
Closing the HPC-Cloud Convergence Gap: Multi-Tenant Slingshot RDMA for Kubernetes
by: Friese, Philipp A., et al.
Published: (2025)
by: Friese, Philipp A., et al.
Published: (2025)
SIREN: Software Identification and Recognition in HPC Systems
by: Jakobsche, Thomas, et al.
Published: (2025)
by: Jakobsche, Thomas, et al.
Published: (2025)
Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters
by: Jain, Rutwik, et al.
Published: (2026)
by: Jain, Rutwik, et al.
Published: (2026)
HPC with Enhanced User Separation
by: Prout, Andrew, et al.
Published: (2024)
by: Prout, Andrew, et al.
Published: (2024)
Analysis of the carbon footprint of HPC
by: Benhari, Abdessalam, et al.
Published: (2025)
by: Benhari, Abdessalam, et al.
Published: (2025)
Bridging Paradigms: Designing for HPC-Quantum Convergence
by: Shehata, Amir, et al.
Published: (2025)
by: Shehata, Amir, et al.
Published: (2025)
Towards an Adaptive Runtime System for Cloud-Native HPC
by: Bhosale, Aditya, et al.
Published: (2026)
by: Bhosale, Aditya, et al.
Published: (2026)
An Autonomy Loop for Dynamic HPC Job Time Limit Adjustment
by: Jakobsche, Thomas, et al.
Published: (2025)
by: Jakobsche, Thomas, et al.
Published: (2025)
MRSch: Multi-Resource Scheduling for HPC
by: Li, Boyang, et al.
Published: (2024)
by: Li, Boyang, et al.
Published: (2024)
Report on Challenges of Practical Reproducibility for Systems and HPC Computer Science
by: Keahey, Kate, et al.
Published: (2025)
by: Keahey, Kate, et al.
Published: (2025)
Sarus Suite: Cloud-native Containers for HPC
by: Madonna, Alberto, et al.
Published: (2026)
by: Madonna, Alberto, et al.
Published: (2026)
Energy-aware operation of HPC systems in Germany
by: Suarez, Estela, et al.
Published: (2024)
by: Suarez, Estela, et al.
Published: (2024)
Hierarchical Resource Partitioning on Modern GPUs: A Reinforcement Learning Approach
by: Saroliya, Urvij, et al.
Published: (2024)
by: Saroliya, Urvij, et al.
Published: (2024)
UNR: Unified Notifiable RMA Library for HPC
by: Feng, Guangnan, et al.
Published: (2024)
by: Feng, Guangnan, et al.
Published: (2024)
Wilkins: HPC In Situ Workflows Made Easy
by: Yildiz, Orcun, et al.
Published: (2024)
by: Yildiz, Orcun, et al.
Published: (2024)
A HPC Co-Scheduler with Reinforcement Learning
by: Souza, Abel, et al.
Published: (2024)
by: Souza, Abel, et al.
Published: (2024)
Software Resource Disaggregation for HPC with Serverless Computing
by: Copik, Marcin, et al.
Published: (2024)
by: Copik, Marcin, et al.
Published: (2024)
An Elastic Job Scheduler for HPC Applications on the Cloud
by: Bhosale, Aditya, et al.
Published: (2025)
by: Bhosale, Aditya, et al.
Published: (2025)
Characterizing the Impact of Congestion in Modern HPC Interconnects
by: Piarulli, Lorenzo, et al.
Published: (2026)
by: Piarulli, Lorenzo, et al.
Published: (2026)
Towards Energy Efficient Co-Scheduling in HPC
by: Zheng, Zhong, et al.
Published: (2026)
by: Zheng, Zhong, et al.
Published: (2026)
Leveraging Teaching on Demand: Approaching HPC to Undergrads
by: Catalán, S., et al.
Published: (2026)
by: Catalán, S., et al.
Published: (2026)
Beyond Pre-Training: The Full Lifecycle of Foundation Models on HPC Systems
by: Conciatore, Dino, et al.
Published: (2026)
by: Conciatore, Dino, et al.
Published: (2026)
Wave-Based Dispatch for Circuit Cutting in Hybrid HPC--Quantum Systems
by: García-Raigada, Ricard S., et al.
Published: (2026)
by: García-Raigada, Ricard S., et al.
Published: (2026)
MPI-over-CXL: Enhancing Communication Efficiency in Distributed HPC Systems
by: Kwon, Miryeong, et al.
Published: (2025)
by: Kwon, Miryeong, et al.
Published: (2025)
A Performance Analysis of Task Scheduling for UQ Workflows on HPC Systems
by: Loi, Chung Ming, et al.
Published: (2025)
by: Loi, Chung Ming, et al.
Published: (2025)
Scalable HPC Job Scheduling and Resource Management in SST
by: Abdurahman, Abubeker, et al.
Published: (2025)
by: Abdurahman, Abubeker, et al.
Published: (2025)
LLM as HPC Expert: Extending RAG Architecture for HPC Data
by: Miyashita, Yusuke, et al.
Published: (2024)
by: Miyashita, Yusuke, et al.
Published: (2024)
Incisor: Ex Ante Cloud Instance Selection for HPC Jobs
by: Laurenzano, Michael A., et al.
Published: (2026)
by: Laurenzano, Michael A., et al.
Published: (2026)
Modeling the Carbon Footprint of HPC: The Top 500 and EasyC
by: Rao, Varsha, et al.
Published: (2025)
by: Rao, Varsha, et al.
Published: (2025)
Similar Items
-
Orchestrated Co-scheduling, Resource Partitioning, and Power Capping on CPU-GPU Heterogeneous Systems via Machine Learning
by: Saba, Issa, et al.
Published: (2024) -
Resource Optimization with MPI Process Malleability for Dynamic Workloads in HPC Clusters
by: Iserte, Sergio, et al.
Published: (2025) -
MPI Malleability Validation under Replayed Real-World HPC Conditions
by: Iserte, S., et al.
Published: (2026) -
Optimizing Hardware Resource Partitioning and Job Allocations on Modern GPUs under Power Caps
by: Arima, Eishi, et al.
Published: (2024) -
Evaluating Malleable Job Scheduling in HPC Clusters using Real-World Workloads
by: Zojer, Patrick, et al.
Published: (2026)