DMRlib: Easy-coding and Efficient Resource Management for Job Malleability
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Iserte, Sergio, Mayo, Rafael, Quintana-Ortí, Enrique S., Peña, Antonio J. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Malleable Molecular Dynamics Simulations with GROMACS and DMR
von: Sandås, Petter, et al.
Veröffentlicht: (2026)
von: Sandås, Petter, et al.
Veröffentlicht: (2026)
Resource Optimization with MPI Process Malleability for Dynamic Workloads in HPC Clusters
von: Iserte, Sergio, et al.
Veröffentlicht: (2025)
von: Iserte, Sergio, et al.
Veröffentlicht: (2025)
MPI Malleability Validation under Replayed Real-World HPC Conditions
von: Iserte, S., et al.
Veröffentlicht: (2026)
von: Iserte, S., et al.
Veröffentlicht: (2026)
A Test Taxonomy and Continuous Integration Ecosystem for Dynamic Resource Management in HPC
von: Sandås, Petter, et al.
Veröffentlicht: (2026)
von: Sandås, Petter, et al.
Veröffentlicht: (2026)
Mapping Parallel Matrix Multiplication in GotoBLAS2 to the AMD Versal ACAP for Deep Learning
von: Lei, Jie, et al.
Veröffentlicht: (2024)
von: Lei, Jie, et al.
Veröffentlicht: (2024)
Towards the Democratization and Standardization of Dynamic Resources with MPI Spawning
von: Iserte, Sergio, et al.
Veröffentlicht: (2026)
von: Iserte, Sergio, et al.
Veröffentlicht: (2026)
Evaluating Malleable Job Scheduling in HPC Clusters using Real-World Workloads
von: Zojer, Patrick, et al.
Veröffentlicht: (2026)
von: Zojer, Patrick, et al.
Veröffentlicht: (2026)
Energy-Efficient Real-Time Job Mapping and Resource Management in Mobile-Edge Computing
von: Gao, Chuanchao, et al.
Veröffentlicht: (2025)
von: Gao, Chuanchao, et al.
Veröffentlicht: (2025)
Wave-Based Dispatch for Circuit Cutting in Hybrid HPC--Quantum Systems
von: García-Raigada, Ricard S., et al.
Veröffentlicht: (2026)
von: García-Raigada, Ricard S., et al.
Veröffentlicht: (2026)
Scalable HPC Job Scheduling and Resource Management in SST
von: Abdurahman, Abubeker, et al.
Veröffentlicht: (2025)
von: Abdurahman, Abubeker, et al.
Veröffentlicht: (2025)
A Study on the Performance of Distributed Training of Data-driven CFD Simulations
von: Iserte, Sergio, et al.
Veröffentlicht: (2026)
von: Iserte, Sergio, et al.
Veröffentlicht: (2026)
Leveraging Teaching on Demand: Approaching HPC to Undergrads
von: Catalán, S., et al.
Veröffentlicht: (2026)
von: Catalán, S., et al.
Veröffentlicht: (2026)
Mean field optimal Core Allocation across Malleable jobs
von: Li, Zhouzi, et al.
Veröffentlicht: (2026)
von: Li, Zhouzi, et al.
Veröffentlicht: (2026)
Fast Truncated SVD of Sparse and Dense Matrices on Graphics Processors
von: Tomas, Andres E., et al.
Veröffentlicht: (2024)
von: Tomas, Andres E., et al.
Veröffentlicht: (2024)
Flora: Efficient Cloud Resource Selection for Big Data Processing via Job Classification
von: Will, Jonathan, et al.
Veröffentlicht: (2025)
von: Will, Jonathan, et al.
Veröffentlicht: (2025)
Venn: Resource Management for Collaborative Learning Jobs
von: Liu, Jiachen, et al.
Veröffentlicht: (2023)
von: Liu, Jiachen, et al.
Veröffentlicht: (2023)
On the Convergence of Malleability and the HPC PowerStack: Exploiting Dynamism in Over-Provisioned and Power-Constrained HPC Systems
von: Arima, Eishi, et al.
Veröffentlicht: (2024)
von: Arima, Eishi, et al.
Veröffentlicht: (2024)
PSI/J: A Portable Interface for Submitting, Monitoring, and Managing Jobs
von: Hategan-Marandiuc, Mihael, et al.
Veröffentlicht: (2023)
von: Hategan-Marandiuc, Mihael, et al.
Veröffentlicht: (2023)
Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization
von: Li, Haoyang, et al.
Veröffentlicht: (2024)
von: Li, Haoyang, et al.
Veröffentlicht: (2024)
Resilient Auto-Scaling of Microservice Architectures with Efficient Resource Management
von: Ahmad, Hussain, et al.
Veröffentlicht: (2025)
von: Ahmad, Hussain, et al.
Veröffentlicht: (2025)
LOw-cOst yet High-Performant Sparse Matrix-Matrix Multiplication on Arm SME Architectures
von: Lei, Kelun, et al.
Veröffentlicht: (2025)
von: Lei, Kelun, et al.
Veröffentlicht: (2025)
Communication-and-Computation Efficient Split Federated Learning: Gradient Aggregation and Resource Management
von: Liang, Yipeng, et al.
Veröffentlicht: (2025)
von: Liang, Yipeng, et al.
Veröffentlicht: (2025)
Optimizing Hardware Resource Partitioning and Job Allocations on Modern GPUs under Power Caps
von: Arima, Eishi, et al.
Veröffentlicht: (2024)
von: Arima, Eishi, et al.
Veröffentlicht: (2024)
FaaS Is Not Enough: Serverless Handling of Burst-Parallel Jobs
von: Barcelona-Pons, Daniel, et al.
Veröffentlicht: (2024)
von: Barcelona-Pons, Daniel, et al.
Veröffentlicht: (2024)
Scheduling Deep Learning Jobs in Multi-Tenant GPU Clusters via Wise Resource Sharing
von: Luo, Yizhou, et al.
Veröffentlicht: (2024)
von: Luo, Yizhou, et al.
Veröffentlicht: (2024)
COUNTER: Cluster GCN based Energy Efficient Resource Management for Sustainable Cloud Computing Environments
von: Wang, Han, et al.
Veröffentlicht: (2025)
von: Wang, Han, et al.
Veröffentlicht: (2025)
The Fused Kernel Library: A C++ API to Develop Highly-Efficient GPU Libraries
von: Amoros, Oscar, et al.
Veröffentlicht: (2025)
von: Amoros, Oscar, et al.
Veröffentlicht: (2025)
Wilkins: HPC In Situ Workflows Made Easy
von: Yildiz, Orcun, et al.
Veröffentlicht: (2024)
von: Yildiz, Orcun, et al.
Veröffentlicht: (2024)
Metronome: Efficient Scheduling for Periodic Traffic Jobs with Network and Priority Awareness
von: Jiang, Hao, et al.
Veröffentlicht: (2025)
von: Jiang, Hao, et al.
Veröffentlicht: (2025)
SPARS: A Reinforcement Learning-Enabled Simulator for Power Management in HPC Job Scheduling
von: Amrizal, Muhammad Alfian, et al.
Veröffentlicht: (2025)
von: Amrizal, Muhammad Alfian, et al.
Veröffentlicht: (2025)
Parallel Reduced Order Modeling for Digital Twins using High-Performance Computing Workflows
von: de Parga, S. Ares, et al.
Veröffentlicht: (2024)
von: de Parga, S. Ares, et al.
Veröffentlicht: (2024)
Driving Computational Efficiency in Large-Scale Platforms using HPC Technologies
von: Mendez, Alexander Martinez, et al.
Veröffentlicht: (2026)
von: Mendez, Alexander Martinez, et al.
Veröffentlicht: (2026)
Modeling the Carbon Footprint of HPC: The Top 500 and EasyC
von: Rao, Varsha, et al.
Veröffentlicht: (2025)
von: Rao, Varsha, et al.
Veröffentlicht: (2025)
Dynamic Solutions for Hybrid Quantum-HPC Resource Allocation
von: Rocco, Roberto, et al.
Veröffentlicht: (2025)
von: Rocco, Roberto, et al.
Veröffentlicht: (2025)
Deep Reinforcement Learning for Job Scheduling and Resource Management in Cloud Computing: An Algorithm-Level Review
von: Gu, Yan, et al.
Veröffentlicht: (2025)
von: Gu, Yan, et al.
Veröffentlicht: (2025)
Dynamic Resource Manager for Automating Deployments in the Computing Continuum
von: Samani, Zahra Najafabadi, et al.
Veröffentlicht: (2024)
von: Samani, Zahra Najafabadi, et al.
Veröffentlicht: (2024)
Embedded Made Easy -- Rethinking Embedded + Cloud Software Development (WIP)
von: Arnold, Anthony, et al.
Veröffentlicht: (2026)
von: Arnold, Anthony, et al.
Veröffentlicht: (2026)
A Conflict-Aware Resource Management Framework for the Computing Continuum
von: Popescu-Vifor, Vlad, et al.
Veröffentlicht: (2025)
von: Popescu-Vifor, Vlad, et al.
Veröffentlicht: (2025)
Decentralized Learning Made Easy with DecentralizePy
von: Dhasade, Akash, et al.
Veröffentlicht: (2023)
von: Dhasade, Akash, et al.
Veröffentlicht: (2023)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Malleable Molecular Dynamics Simulations with GROMACS and DMR
von: Sandås, Petter, et al.
Veröffentlicht: (2026) -
Resource Optimization with MPI Process Malleability for Dynamic Workloads in HPC Clusters
von: Iserte, Sergio, et al.
Veröffentlicht: (2025) -
MPI Malleability Validation under Replayed Real-World HPC Conditions
von: Iserte, S., et al.
Veröffentlicht: (2026) -
A Test Taxonomy and Continuous Integration Ecosystem for Dynamic Resource Management in HPC
von: Sandås, Petter, et al.
Veröffentlicht: (2026) -
Mapping Parallel Matrix Multiplication in GotoBLAS2 to the AMD Versal ACAP for Deep Learning
von: Lei, Jie, et al.
Veröffentlicht: (2024)