Optimizing Checkpoint-Restart Mechanisms for HPC with DMTCP in Containers at NERSC

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Timalsina, Madan, Gerhardt, Lisa, Tyler, Nicholas, Blaschke, Johannes P., Arndt, William
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929438971658240
author Timalsina, Madan
Gerhardt, Lisa
Tyler, Nicholas
Blaschke, Johannes P.
Arndt, William
author_facet Timalsina, Madan
Gerhardt, Lisa
Tyler, Nicholas
Blaschke, Johannes P.
Arndt, William
contents This paper presents an in-depth examination of checkpoint-restart mechanisms in High-Performance Computing (HPC). It focuses on the use of Distributed MultiThreaded CheckPointing (DMTCP) in various computational settings, including both within and outside of containers. The study is grounded in real-world applications running on NERSC Perlmutter, a state-of-the-art supercomputing system. We discuss the advantages of checkpoint-restart (C/R) in managing complex and lengthy computations in HPC, highlighting its efficiency and reliability in such environments. The role of DMTCP in enhancing these workflows, especially in multi-threaded and distributed applications, is thoroughly explored. Additionally, the paper delves into the use of HPC containers, such as Shifter and Podman-HPC, which aid in the management of computational tasks, ensuring uniform performance across different environments. The methods, results, and potential future directions of this research, including its application in various scientific domains, are also covered, showcasing the critical advancements made in computational methodologies through this study.
format Preprint
id arxiv_https___arxiv_org_abs_2407_19117
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Optimizing Checkpoint-Restart Mechanisms for HPC with DMTCP in Containers at NERSC
Timalsina, Madan
Gerhardt, Lisa
Tyler, Nicholas
Blaschke, Johannes P.
Arndt, William
Distributed, Parallel, and Cluster Computing
Software Engineering
This paper presents an in-depth examination of checkpoint-restart mechanisms in High-Performance Computing (HPC). It focuses on the use of Distributed MultiThreaded CheckPointing (DMTCP) in various computational settings, including both within and outside of containers. The study is grounded in real-world applications running on NERSC Perlmutter, a state-of-the-art supercomputing system. We discuss the advantages of checkpoint-restart (C/R) in managing complex and lengthy computations in HPC, highlighting its efficiency and reliability in such environments. The role of DMTCP in enhancing these workflows, especially in multi-threaded and distributed applications, is thoroughly explored. Additionally, the paper delves into the use of HPC containers, such as Shifter and Podman-HPC, which aid in the management of computational tasks, ensuring uniform performance across different environments. The methods, results, and potential future directions of this research, including its application in various scientific domains, are also covered, showcasing the critical advancements made in computational methodologies through this study.
title Optimizing Checkpoint-Restart Mechanisms for HPC with DMTCP in Containers at NERSC
topic Distributed, Parallel, and Cluster Computing
Software Engineering
url https://arxiv.org/abs/2407.19117