Optimizing Checkpoint-Restart Mechanisms for HPC with DMTCP in Containers at NERSC
Fuente:
arXiv
Saved in:
| Main Authors: | Timalsina, Madan, Gerhardt, Lisa, Tyler, Nicholas, Blaschke, Johannes P., Arndt, William |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Analysis of HPC and Edge Architectures in the Cloud
by: Santillan, Steven, et al.
Published: (2025)
by: Santillan, Steven, et al.
Published: (2025)
Addressing Reproducibility Challenges in HPC with Continuous Integration
by: Hayot-Sasson, Valérie, et al.
Published: (2025)
by: Hayot-Sasson, Valérie, et al.
Published: (2025)
A Test Taxonomy and Continuous Integration Ecosystem for Dynamic Resource Management in HPC
by: Sandås, Petter, et al.
Published: (2026)
by: Sandås, Petter, et al.
Published: (2026)
Hydra: Brokering Cloud and HPC Resources to Support the Execution of Heterogeneous Workloads at Scale
by: Alsaadi, Aymen, et al.
Published: (2024)
by: Alsaadi, Aymen, et al.
Published: (2024)
LLM-HPC++: Evaluating LLM-Generated Modern C++ and MPI+OpenMP Codes for Scalable Mandelbrot Set Computation
by: Diehl, Patrick, et al.
Published: (2025)
by: Diehl, Patrick, et al.
Published: (2025)
Container-level Energy Observability in Kubernetes Clusters
by: Pijnacker, Bjorn, et al.
Published: (2025)
by: Pijnacker, Bjorn, et al.
Published: (2025)
LLMs as Packagers of HPC Software
by: Melone, Caetano, et al.
Published: (2025)
by: Melone, Caetano, et al.
Published: (2025)
VibeCodeHPC: An Agent-Based Iterative Prompting Auto-Tuner for HPC Code Generation Using LLMs
by: Hayashi, Shun-ichiro, et al.
Published: (2025)
by: Hayashi, Shun-ichiro, et al.
Published: (2025)
From Edge to HPC: Investigating Cross-Facility Data Streaming Architectures
by: George, Anjus, et al.
Published: (2025)
by: George, Anjus, et al.
Published: (2025)
Do Large Language Models Understand Performance Optimization?
by: Cui, Bowen, et al.
Published: (2025)
by: Cui, Bowen, et al.
Published: (2025)
Towards an Optimized Benchmarking Platform for CI/CD Pipelines
by: Japke, Nils, et al.
Published: (2025)
by: Japke, Nils, et al.
Published: (2025)
SPES: Towards Optimizing Performance-Resource Trade-Off for Serverless Functions
by: Lee, Cheryl, et al.
Published: (2024)
by: Lee, Cheryl, et al.
Published: (2024)
Comprehensive Review of Performance Optimization Strategies for Serverless Applications on AWS Lambda
by: Bechir, Mohamed Lemine El, et al.
Published: (2024)
by: Bechir, Mohamed Lemine El, et al.
Published: (2024)
CloudHeatMap: Heatmap-Based Monitoring for Large-Scale Cloud Systems
by: Sohana, Sarah, et al.
Published: (2024)
by: Sohana, Sarah, et al.
Published: (2024)
HPC-Coder-V2: Studying Code LLMs Across Low-Resource Parallel Languages
by: Chaturvedi, Aman, et al.
Published: (2024)
by: Chaturvedi, Aman, et al.
Published: (2024)
HPCAgentTester: A Multi-Agent LLM Approach for Enhanced HPC Unit Test Generation
by: Karanjai, Rabimba, et al.
Published: (2025)
by: Karanjai, Rabimba, et al.
Published: (2025)
A Reference Architecture for Governance of Cloud Native Applications
by: Pourmajidi, William, et al.
Published: (2023)
by: Pourmajidi, William, et al.
Published: (2023)
Leveraging AI for Productive and Trustworthy HPC Software: Challenges and Research Directions
by: Teranishi, Keita, et al.
Published: (2025)
by: Teranishi, Keita, et al.
Published: (2025)
Supercharging Federated Learning with Flower and NVIDIA FLARE
by: Roth, Holger R., et al.
Published: (2024)
by: Roth, Holger R., et al.
Published: (2024)
Securing Confidential Data For Distributed Software Development Teams: Encrypted Container File
by: Bauer, Tobias J., et al.
Published: (2024)
by: Bauer, Tobias J., et al.
Published: (2024)
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
by: Nichols, Daniel, et al.
Published: (2025)
by: Nichols, Daniel, et al.
Published: (2025)
Checkpoint and Restart: An Energy Consumption Characterization in Clusters
by: Moran, Marina, et al.
Published: (2024)
by: Moran, Marina, et al.
Published: (2024)
Encrypted Container File: Design and Implementation of a Hybrid-Encrypted Multi-Recipient File Structure
by: Bauer, Tobias J., et al.
Published: (2024)
by: Bauer, Tobias J., et al.
Published: (2024)
SeBS-Flow: Benchmarking Serverless Cloud Function Workflows
by: Schmid, Larissa, et al.
Published: (2024)
by: Schmid, Larissa, et al.
Published: (2024)
A Unifying Framework to Enable Artificial Intelligence in High Performance Computing Workflows
by: Domke, Jens, et al.
Published: (2025)
by: Domke, Jens, et al.
Published: (2025)
$μ$OpTime: Statically Reducing the Execution Time of Microbenchmark Suites Using Stability Metrics
by: Japke, Nils, et al.
Published: (2025)
by: Japke, Nils, et al.
Published: (2025)
Adaptable TeaStore
by: Bliudze, Simon, et al.
Published: (2024)
by: Bliudze, Simon, et al.
Published: (2024)
Building Castles in the Cloud: Architecting Resilient and Scalable Infrastructure
by: Gundla, Naresh Kumar
Published: (2024)
by: Gundla, Naresh Kumar
Published: (2024)
Histrio: a Serverless Actor System
by: Buttiglieri, Giorgio Natale, et al.
Published: (2024)
by: Buttiglieri, Giorgio Natale, et al.
Published: (2024)
GitFarm: Git as a Service for Large-Scale Monorepos
by: Dwivedi, Preetam, et al.
Published: (2026)
by: Dwivedi, Preetam, et al.
Published: (2026)
Predictive Autoscaling for Node.js on Kubernetes: Lower Latency, Right-Sized Capacity
by: Tymoshenko, Ivan, et al.
Published: (2026)
by: Tymoshenko, Ivan, et al.
Published: (2026)
Efficiently Reproducing Distributed Workflows in Notebook-based Systems
by: Azaz, Talha, et al.
Published: (2026)
by: Azaz, Talha, et al.
Published: (2026)
AdaptiFlow: An Extensible Framework for Event-Driven Autonomy in Cloud Microservices
by: Ndadji, Brice Arléon Zemtsop, et al.
Published: (2025)
by: Ndadji, Brice Arléon Zemtsop, et al.
Published: (2025)
AlertGuardian: Intelligent Alert Life-Cycle Management for Large-scale Cloud Systems
by: Yu, Guangba, et al.
Published: (2026)
by: Yu, Guangba, et al.
Published: (2026)
Carbon-aware Software Services
by: Forti, Stefano, et al.
Published: (2024)
by: Forti, Stefano, et al.
Published: (2024)
CARISMA: CAR-Integrated Service Mesh Architecture
by: Klein, Kevin, et al.
Published: (2024)
by: Klein, Kevin, et al.
Published: (2024)
Umbilical Choir: Automated Live Testing for Edge-To-Cloud FaaS Applications
by: Malekabbasi, Mohammadreza, et al.
Published: (2025)
by: Malekabbasi, Mohammadreza, et al.
Published: (2025)
Specx: a C++ task-based runtime system for heterogeneous distributed architectures
by: Cardosi, Paul, et al.
Published: (2023)
by: Cardosi, Paul, et al.
Published: (2023)
ATOM: Asynchronous Training of Massive Models for Deep Learning in a Decentralized Environment
by: Wu, Xiaofeng, et al.
Published: (2024)
by: Wu, Xiaofeng, et al.
Published: (2024)
FlowUnits: Extending Dataflow for the Edge-to-Cloud Computing Continuum
by: Chini, Fabio, et al.
Published: (2025)
by: Chini, Fabio, et al.
Published: (2025)
Similar Items
-
An Analysis of HPC and Edge Architectures in the Cloud
by: Santillan, Steven, et al.
Published: (2025) -
Addressing Reproducibility Challenges in HPC with Continuous Integration
by: Hayot-Sasson, Valérie, et al.
Published: (2025) -
A Test Taxonomy and Continuous Integration Ecosystem for Dynamic Resource Management in HPC
by: Sandås, Petter, et al.
Published: (2026) -
Hydra: Brokering Cloud and HPC Resources to Support the Execution of Heterogeneous Workloads at Scale
by: Alsaadi, Aymen, et al.
Published: (2024) -
LLM-HPC++: Evaluating LLM-Generated Modern C++ and MPI+OpenMP Codes for Scalable Mandelbrot Set Computation
by: Diehl, Patrick, et al.
Published: (2025)