Enabling Practical Transparent Checkpointing for MPI: A Topological Sort Approach
Fuente:
arXiv
Guardado en:
| Autores principales: | Xu, Yao, Cooperman, Gene |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The Case for ABI Interoperability in a Fault Tolerant MPI
por: Xu, Yao, et al.
Publicado: (2025)
por: Xu, Yao, et al.
Publicado: (2025)
DNA sequence alignment: An assignment for OpenMP, MPI, and CUDA/OpenCL
por: Gonzalez-Escribano, Arturo, et al.
Publicado: (2024)
por: Gonzalez-Escribano, Arturo, et al.
Publicado: (2024)
AutoTSMM: An Auto-tuning Framework for Building High-Performance Tall-and-Skinny Matrix-Matrix Multiplication on CPUs
por: Li, Chendi, et al.
Publicado: (2022)
por: Li, Chendi, et al.
Publicado: (2022)
Optimizing Fine-Grained Parallelism Through Dynamic Load Balancing on Multi-Socket Many-Core Systems
por: Wang, Wenyi, et al.
Publicado: (2025)
por: Wang, Wenyi, et al.
Publicado: (2025)
VSS Challenge Problem: Verifying the Correctness of AllReduce Algorithms in the MPICH Implementation of MPI
por: Hovland, Paul D.
Publicado: (2025)
por: Hovland, Paul D.
Publicado: (2025)
Challenging Portability Paradigms: FPGA Acceleration Using SYCL and OpenCL
por: de Castro, Manuel, et al.
Publicado: (2024)
por: de Castro, Manuel, et al.
Publicado: (2024)
A C++17 Thread Pool for High-Performance Scientific Computing
por: Shoshany, Barak
Publicado: (2021)
por: Shoshany, Barak
Publicado: (2021)
Stream parallel skeleton optimization
por: Aldinucci, Marco, et al.
Publicado: (2024)
por: Aldinucci, Marco, et al.
Publicado: (2024)
StreamFlow: cross-breeding cloud with HPC
por: Colonnelli, Iacopo, et al.
Publicado: (2020)
por: Colonnelli, Iacopo, et al.
Publicado: (2020)
HotSwap: Enabling Live Dependency Sharing in Serverless Computing
por: Li, Rui, et al.
Publicado: (2024)
por: Li, Rui, et al.
Publicado: (2024)
Sharded Elimination and Combining for Highly-Efficient Concurrent Stacks
por: Singh, Ajay, et al.
Publicado: (2026)
por: Singh, Ajay, et al.
Publicado: (2026)
VeriFx: Correct Replicated Data Types for the Masses
por: De Porre, Kevin, et al.
Publicado: (2022)
por: De Porre, Kevin, et al.
Publicado: (2022)
Dynamic Memory Management on GPUs with SYCL
por: Standish, Russell K.
Publicado: (2025)
por: Standish, Russell K.
Publicado: (2025)
Hybrid Quantum-HPC Middleware Systems for Adaptive Resource, Workload and Task Management
por: Mantha, Pradeep, et al.
Publicado: (2026)
por: Mantha, Pradeep, et al.
Publicado: (2026)
NVLang: Unified Static Typing for Actor-Based Concurrency on the BEAM
por: Guerreiro, Miguel de Oliveira
Publicado: (2025)
por: Guerreiro, Miguel de Oliveira
Publicado: (2025)
Utilizing Sparsity in the GPU-accelerated Assembly of Schur Complement Matrices in Domain Decomposition Methods
por: Homola, Jakub, et al.
Publicado: (2025)
por: Homola, Jakub, et al.
Publicado: (2025)
Joint Training on AMD and NVIDIA GPUs
por: Hu, Jon, et al.
Publicado: (2026)
por: Hu, Jon, et al.
Publicado: (2026)
Fancy Some Chips for Your TeaStore? Modeling the Control of an Adaptable Discrete System
por: Gallone, Anna, et al.
Publicado: (2025)
por: Gallone, Anna, et al.
Publicado: (2025)
How to Relax Instantly: Elastic Relaxation of Concurrent Data Structures
por: von Geijer, Kåre, et al.
Publicado: (2024)
por: von Geijer, Kåre, et al.
Publicado: (2024)
SpaDA: A Spatial Dataflow Architecture Programming Language
por: Gianinazzi, Lukas, et al.
Publicado: (2025)
por: Gianinazzi, Lukas, et al.
Publicado: (2025)
Minimum Cost Loop Nests for Contraction of a Sparse Tensor with a Tensor Network
por: Kanakagiri, Raghavendra, et al.
Publicado: (2023)
por: Kanakagiri, Raghavendra, et al.
Publicado: (2023)
Categorical Message Passing Language (CaMPL) for programmers
por: Hashimoto, Daniel Kiyoshi, et al.
Publicado: (2026)
por: Hashimoto, Daniel Kiyoshi, et al.
Publicado: (2026)
An Evaluation of Massively Parallel Algorithms for DFA Minimization
por: Martens, Jan, et al.
Publicado: (2024)
por: Martens, Jan, et al.
Publicado: (2024)
NM-SpMM: Accelerating Matrix Multiplication Using N:M Sparsity with GPGPU
por: Ma, Cong, et al.
Publicado: (2025)
por: Ma, Cong, et al.
Publicado: (2025)
Introducing SWIRL: An Intermediate Representation Language for Scientific Workflows
por: Colonnelli, Iacopo, et al.
Publicado: (2024)
por: Colonnelli, Iacopo, et al.
Publicado: (2024)
MATCH: Model-Aware TVM-based Compilation for Heterogeneous Edge Devices
por: Hamdi, Mohamed Amine, et al.
Publicado: (2024)
por: Hamdi, Mohamed Amine, et al.
Publicado: (2024)
Accelerating Gravitational $N$-Body Simulations Using the RISC-V-Based Tenstorrent Wormhole
por: Almerol, Jenny Lynn, et al.
Publicado: (2025)
por: Almerol, Jenny Lynn, et al.
Publicado: (2025)
Static Batching of Irregular Workloads on GPUs: Framework and Application to Efficient MoE Model Inference
por: Li, Yinghan, et al.
Publicado: (2025)
por: Li, Yinghan, et al.
Publicado: (2025)
Scheduler-Driven Job Atomization
por: Konopa, Michal, et al.
Publicado: (2025)
por: Konopa, Michal, et al.
Publicado: (2025)
JASDA: Introducing Job-Aware Scheduling in Scheduler-Driven Job Atomization
por: Konopa, Michal, et al.
Publicado: (2025)
por: Konopa, Michal, et al.
Publicado: (2025)
Vectorized Adaptive Histograms for Sparse Oblique Forests
por: Lubonja, Ariel, et al.
Publicado: (2026)
por: Lubonja, Ariel, et al.
Publicado: (2026)
CUNQA: a Distributed Quantum Computing emulator for HPC
por: Vázquez-Pérez, Jorge, et al.
Publicado: (2025)
por: Vázquez-Pérez, Jorge, et al.
Publicado: (2025)
pLUTo: Enabling Massively Parallel Computation in DRAM via Lookup Tables
por: Ferreira, João Dinis, et al.
Publicado: (2021)
por: Ferreira, João Dinis, et al.
Publicado: (2021)
HierarchicalKV: A GPU Hash Table with Cache Semantics for Continuous Online Embedding Storage
por: Rong, Haidong, et al.
Publicado: (2026)
por: Rong, Haidong, et al.
Publicado: (2026)
CRIUgpu: Transparent Checkpointing of GPU-Accelerated Workloads
por: Stoyanov, Radostin, et al.
Publicado: (2025)
por: Stoyanov, Radostin, et al.
Publicado: (2025)
Hector: An Efficient Programming and Compilation Framework for Implementing Relational Graph Neural Networks in GPU Architectures
por: Wu, Kun, et al.
Publicado: (2023)
por: Wu, Kun, et al.
Publicado: (2023)
Construction of a Byzantine Linearizable SWMR Atomic Register from SWSR Atomic Registers
por: Kshemkalyani, Ajay D., et al.
Publicado: (2024)
por: Kshemkalyani, Ajay D., et al.
Publicado: (2024)
Breaking (Global) Barriers in Parallel Stochastic Optimization with Wait-Avoiding Group Averaging
por: Li, Shigang, et al.
Publicado: (2020)
por: Li, Shigang, et al.
Publicado: (2020)
Improving inference time in multi-TPU systems with profiled model segmentation
por: Villarrubia, Jorge, et al.
Publicado: (2025)
por: Villarrubia, Jorge, et al.
Publicado: (2025)
Balanced segmentation of CNNs for multi-TPU inference
por: Villarrubia, Jorge, et al.
Publicado: (2025)
por: Villarrubia, Jorge, et al.
Publicado: (2025)
Ejemplares similares
-
The Case for ABI Interoperability in a Fault Tolerant MPI
por: Xu, Yao, et al.
Publicado: (2025) -
DNA sequence alignment: An assignment for OpenMP, MPI, and CUDA/OpenCL
por: Gonzalez-Escribano, Arturo, et al.
Publicado: (2024) -
AutoTSMM: An Auto-tuning Framework for Building High-Performance Tall-and-Skinny Matrix-Matrix Multiplication on CPUs
por: Li, Chendi, et al.
Publicado: (2022) -
Optimizing Fine-Grained Parallelism Through Dynamic Load Balancing on Multi-Socket Many-Core Systems
por: Wang, Wenyi, et al.
Publicado: (2025) -
VSS Challenge Problem: Verifying the Correctness of AllReduce Algorithms in the MPICH Implementation of MPI
por: Hovland, Paul D.
Publicado: (2025)