Universal Checkpointing: A Flexible and Efficient Distributed Checkpointing System for Large-Scale DNN Training with Reconfigurable Parallelis
Fuente:
arXiv
Saved in:
| Main Authors: | Lian, Xinyu, Jacobs, Sam Ade, Kurilenko, Lev, Tanaka, Masahiro, Bekman, Stas, Ruwase, Olatunji, Zhang, Minjia |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
by: Lian, Xinyu, et al.
Published: (2025)
by: Lian, Xinyu, et al.
Published: (2025)
Training Ultra Long Context Language Model with Fully Pipelined Distributed Transformer
by: Yao, Jinghan, et al.
Published: (2024)
by: Yao, Jinghan, et al.
Published: (2024)
DeepCompile: A Compiler-Driven Approach to Optimizing Distributed Deep Learning Training
by: Tanaka, Masahiro, et al.
Published: (2025)
by: Tanaka, Masahiro, et al.
Published: (2025)
Cross-Layer Energy Analysis of Multimodal Training on Grace Hopper Superchips
by: Ahmed, Mahmoud, et al.
Published: (2026)
by: Ahmed, Mahmoud, et al.
Published: (2026)
FastPersist: Accelerating Model Checkpointing in Deep Learning
by: Wang, Guanhua, et al.
Published: (2024)
by: Wang, Guanhua, et al.
Published: (2024)
AutoSP: Unlocking Long-Context LLM Training Via Compiler-Based Sequence Parallelism
by: Gupta, Ahan, et al.
Published: (2026)
by: Gupta, Ahan, et al.
Published: (2026)
Optimizing Frequent Checkpointing via Low-Cost Differential for Distributed Training Systems
by: Yao, Chenxuan, et al.
Published: (2025)
by: Yao, Chenxuan, et al.
Published: (2025)
Sparse Checkpointing for Fast and Reliable MoE Training
by: Gandhi, Swapnil, et al.
Published: (2024)
by: Gandhi, Swapnil, et al.
Published: (2024)
Orbax: Distributed Checkpointing with JAX
by: Gaffney, Colin, et al.
Published: (2026)
by: Gaffney, Colin, et al.
Published: (2026)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
by: Wang, Yuxin, et al.
Published: (2023)
by: Wang, Yuxin, et al.
Published: (2023)
Asynchronous Checkpoint for Eventually Consistent Databases
by: Ravishankar, Raaghav, et al.
Published: (2025)
by: Ravishankar, Raaghav, et al.
Published: (2025)
Scrutinizing Variables for Checkpoint Using Automatic Differentiation
by: Huang, Xin, et al.
Published: (2026)
by: Huang, Xin, et al.
Published: (2026)
CRIUgpu: Transparent Checkpointing of GPU-Accelerated Workloads
by: Stoyanov, Radostin, et al.
Published: (2025)
by: Stoyanov, Radostin, et al.
Published: (2025)
Optimal Checkpoint Interval with Availability as an Objective Function
by: Saxena, Nirmal Raj, et al.
Published: (2024)
by: Saxena, Nirmal Raj, et al.
Published: (2024)
Checkpoint and Restart: An Energy Consumption Characterization in Clusters
by: Moran, Marina, et al.
Published: (2024)
by: Moran, Marina, et al.
Published: (2024)
Parcae: Proactive, Liveput-Optimized DNN Training on Preemptible Instances
by: Duan, Jiangfei, et al.
Published: (2024)
by: Duan, Jiangfei, et al.
Published: (2024)
Efficient LLM Inference with Activation Checkpointing and Hybrid Caching
by: Lee, Sanghyeon, et al.
Published: (2025)
by: Lee, Sanghyeon, et al.
Published: (2025)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
by: Zhang, WenZheng, et al.
Published: (2024)
by: Zhang, WenZheng, et al.
Published: (2024)
CRIU -- Checkpoint Restore in Userspace for computational simulations and scientific applications
by: Andrijauskas, Fabio, et al.
Published: (2024)
by: Andrijauskas, Fabio, et al.
Published: (2024)
Understanding LLM Checkpoint/Restore I/O Strategies and Patterns
by: Gossman, Mikaila J., et al.
Published: (2025)
by: Gossman, Mikaila J., et al.
Published: (2025)
Checkmate: Zero-Overhead Model Checkpointing via Network Gradient Replication
by: Bhardwaj, Ankit, et al.
Published: (2025)
by: Bhardwaj, Ankit, et al.
Published: (2025)
Domino: Eliminating Communication in LLM Training via Generic Tensor Slicing and Overlapping
by: Wang, Guanhua, et al.
Published: (2024)
by: Wang, Guanhua, et al.
Published: (2024)
A Flexible Programmable Pipeline Parallelism Framework for Efficient DNN Training
by: Jiang, Lijuan, et al.
Published: (2025)
by: Jiang, Lijuan, et al.
Published: (2025)
MAC-Attention: a Match-Amend-Complete Scheme for Fast and Accurate Attention Computation
by: Yao, Jinghan, et al.
Published: (2026)
by: Yao, Jinghan, et al.
Published: (2026)
CheckMate: Evaluating Checkpointing Protocols for Streaming Dataflows
by: Siachamis, George, et al.
Published: (2024)
by: Siachamis, George, et al.
Published: (2024)
TierCheck: Tiered Checkpointing for Fault Tolerance in Large Language Model Training
by: Han, Shujie, et al.
Published: (2026)
by: Han, Shujie, et al.
Published: (2026)
LLMTailor: A Layer-wise Tailoring Tool for Efficient Checkpointing of Large Language Models
by: Sun, Minqiu, et al.
Published: (2026)
by: Sun, Minqiu, et al.
Published: (2026)
ParaLog: Consistent Host-side Logging for Parallel Checkpoints
by: Chien, Steven W. D., et al.
Published: (2024)
by: Chien, Steven W. D., et al.
Published: (2024)
All is Not Lost: LLM Recovery without Checkpoints
by: Blagoev, Nikolay, et al.
Published: (2025)
by: Blagoev, Nikolay, et al.
Published: (2025)
Towards a Flexible and High-Fidelity Approach to Distributed DNN Training Emulation
by: Liu, Banruo, et al.
Published: (2024)
by: Liu, Banruo, et al.
Published: (2024)
Nezha: Breaking Multi-Rail Network Barriers for Distributed DNN Training
by: Yu, Enda, et al.
Published: (2024)
by: Yu, Enda, et al.
Published: (2024)
Scalable and Adaptively Secure Any-Trust Distributed Key Generation and All-hands Checkpointing
by: Feng, Hanwen, et al.
Published: (2023)
by: Feng, Hanwen, et al.
Published: (2023)
Architectural Foundations for Checkpointing and Restoration in Quantum HPC Systems
by: Guan, Qiang, et al.
Published: (2026)
by: Guan, Qiang, et al.
Published: (2026)
A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO
by: Svedas, Jonas, et al.
Published: (2025)
by: Svedas, Jonas, et al.
Published: (2025)
Efficient N-to-M Checkpointing Algorithm for Finite Element Simulations
by: Ham, David A., et al.
Published: (2024)
by: Ham, David A., et al.
Published: (2024)
Optimizing Checkpoint-Restart Mechanisms for HPC with DMTCP in Containers at NERSC
by: Timalsina, Madan, et al.
Published: (2024)
by: Timalsina, Madan, et al.
Published: (2024)
DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers
by: Maurya, Avinash, et al.
Published: (2026)
by: Maurya, Avinash, et al.
Published: (2026)
DataStates-LLM: Lazy Asynchronous Checkpointing for Large Language Models
by: Maurya, Avinash, et al.
Published: (2024)
by: Maurya, Avinash, et al.
Published: (2024)
MiCRO: Near-Zero Cost Gradient Sparsification for Scaling and Accelerating Distributed DNN Training
by: Yoon, Daegun, et al.
Published: (2023)
by: Yoon, Daegun, et al.
Published: (2023)
PhoenixOS: Concurrent OS-level GPU Checkpoint and Restore with Validated Speculation
by: Wei, Xingda, et al.
Published: (2024)
by: Wei, Xingda, et al.
Published: (2024)
Similar Items
-
SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
by: Lian, Xinyu, et al.
Published: (2025) -
Training Ultra Long Context Language Model with Fully Pipelined Distributed Transformer
by: Yao, Jinghan, et al.
Published: (2024) -
DeepCompile: A Compiler-Driven Approach to Optimizing Distributed Deep Learning Training
by: Tanaka, Masahiro, et al.
Published: (2025) -
Cross-Layer Energy Analysis of Multimodal Training on Grace Hopper Superchips
by: Ahmed, Mahmoud, et al.
Published: (2026) -
FastPersist: Accelerating Model Checkpointing in Deep Learning
by: Wang, Guanhua, et al.
Published: (2024)