Checkmate: Zero-Overhead Model Checkpointing via Network Gradient Replication
Fuente:
arXiv
Saved in:
| Main Authors: | Bhardwaj, Ankit, Wang, Weiyang, Carin, Jeremy, Belay, Adam, Ghobadi, Manya |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KaMPIng: Flexible and (Near) Zero-Overhead C++ Bindings for MPI
by: Uhl, Tim Niklas, et al.
Published: (2024)
by: Uhl, Tim Niklas, et al.
Published: (2024)
Near-Zero-Overhead Freshness for Recommendation Systems via Inference-Side Model Updates
by: Yu, Wenjun, et al.
Published: (2025)
by: Yu, Wenjun, et al.
Published: (2025)
Optimizing Frequent Checkpointing via Low-Cost Differential for Distributed Training Systems
by: Yao, Chenxuan, et al.
Published: (2025)
by: Yao, Chenxuan, et al.
Published: (2025)
Scaling LLM Inference Beyond Amdahl`s Limits via Eliminating Non-Scalable Overheads
by: Zhao, Alan, et al.
Published: (2026)
by: Zhao, Alan, et al.
Published: (2026)
MLTCP: Congestion Control for DNN Training
by: Rajasekaran, Sudarsanan, et al.
Published: (2024)
by: Rajasekaran, Sudarsanan, et al.
Published: (2024)
Asynchronous Checkpoint for Eventually Consistent Databases
by: Ravishankar, Raaghav, et al.
Published: (2025)
by: Ravishankar, Raaghav, et al.
Published: (2025)
Building State Machine Replication Using Practical Network Synchrony
by: Wan, Yiliang, et al.
Published: (2025)
by: Wan, Yiliang, et al.
Published: (2025)
AMSP: Reducing Communication Overhead of ZeRO for Efficient LLM Training
by: Chen, Qiaoling, et al.
Published: (2023)
by: Chen, Qiaoling, et al.
Published: (2023)
Parallelize Over Data Particle Advection: Participation, Ping Pong Particles, and Overhead
by: Wang, Zhe, et al.
Published: (2024)
by: Wang, Zhe, et al.
Published: (2024)
Understanding and Reducing Metadata-Driven Host Overheads in Sampling-Based GNN Training
by: Gong, Yidong, et al.
Published: (2026)
by: Gong, Yidong, et al.
Published: (2026)
Modular Architecture for High-Performance and Low Overhead Data Transfers
by: Swargo, Rasman Mubtasim, et al.
Published: (2025)
by: Swargo, Rasman Mubtasim, et al.
Published: (2025)
CRIUgpu: Transparent Checkpointing of GPU-Accelerated Workloads
by: Stoyanov, Radostin, et al.
Published: (2025)
by: Stoyanov, Radostin, et al.
Published: (2025)
Scrutinizing Variables for Checkpoint Using Automatic Differentiation
by: Huang, Xin, et al.
Published: (2026)
by: Huang, Xin, et al.
Published: (2026)
Optimal Checkpoint Interval with Availability as an Objective Function
by: Saxena, Nirmal Raj, et al.
Published: (2024)
by: Saxena, Nirmal Raj, et al.
Published: (2024)
Checkpoint and Restart: An Energy Consumption Characterization in Clusters
by: Moran, Marina, et al.
Published: (2024)
by: Moran, Marina, et al.
Published: (2024)
Benchmarking Compound AI Applications for Hardware-Software Co-Design
by: Samuthrsindh, Paramuth, et al.
Published: (2026)
by: Samuthrsindh, Paramuth, et al.
Published: (2026)
LLMTailor: A Layer-wise Tailoring Tool for Efficient Checkpointing of Large Language Models
by: Sun, Minqiu, et al.
Published: (2026)
by: Sun, Minqiu, et al.
Published: (2026)
Efficient LLM Inference with Activation Checkpointing and Hybrid Caching
by: Lee, Sanghyeon, et al.
Published: (2025)
by: Lee, Sanghyeon, et al.
Published: (2025)
Sparse Checkpointing for Fast and Reliable MoE Training
by: Gandhi, Swapnil, et al.
Published: (2024)
by: Gandhi, Swapnil, et al.
Published: (2024)
Understanding LLM Checkpoint/Restore I/O Strategies and Patterns
by: Gossman, Mikaila J., et al.
Published: (2025)
by: Gossman, Mikaila J., et al.
Published: (2025)
CRIU -- Checkpoint Restore in Userspace for computational simulations and scientific applications
by: Andrijauskas, Fabio, et al.
Published: (2024)
by: Andrijauskas, Fabio, et al.
Published: (2024)
SafarDB: FPGA-Accelerated Distributed Transactions via Replicated Data Types
by: Saberlatibari, Javad, et al.
Published: (2026)
by: Saberlatibari, Javad, et al.
Published: (2026)
FeedSign: Robust Full-parameter Federated Fine-tuning of Large Models with Extremely Low Communication Overhead of One Bit
by: Cai, Zhijie, et al.
Published: (2025)
by: Cai, Zhijie, et al.
Published: (2025)
Reliable Replication Protocols on SmartNICs
by: Katebzadeh, M. R. Siavash, et al.
Published: (2025)
by: Katebzadeh, M. R. Siavash, et al.
Published: (2025)
Undo and Redo Support for Replicated Registers
by: Stewen, Leo, et al.
Published: (2024)
by: Stewen, Leo, et al.
Published: (2024)
Junctiond: Extending FaaS Runtimes with Kernel-Bypass
by: Saurez, Enrique, et al.
Published: (2024)
by: Saurez, Enrique, et al.
Published: (2024)
A Lightweight Approach for State Machine Replication
by: Cachin, Christian, et al.
Published: (2025)
by: Cachin, Christian, et al.
Published: (2025)
Mangrove: Fast and Parallelizable State Replication for Blockchains
by: Paramonov, Anton, et al.
Published: (2025)
by: Paramonov, Anton, et al.
Published: (2025)
Approaches to Conflict-free Replicated Data Types
by: Almeida, Paulo Sérgio
Published: (2023)
by: Almeida, Paulo Sérgio
Published: (2023)
Linearizability and State-Machine Replication: Is it a match?
by: Hauck, Franz J., et al.
Published: (2024)
by: Hauck, Franz J., et al.
Published: (2024)
The Cost of Garbage Collection for State Machine Replication
by: Liang, Zhiying, et al.
Published: (2024)
by: Liang, Zhiying, et al.
Published: (2024)
A Unified, Practical, and Understandable Model of Non-transactional Consistency Levels in Distributed Replication
by: Hu, Guanzhou, et al.
Published: (2024)
by: Hu, Guanzhou, et al.
Published: (2024)
Ocior: Ultra-Fast Asynchronous Leaderless Consensus with Two-Round Finality, Linear Overhead, and Adaptive Security
by: Chen, Jinyuan
Published: (2025)
by: Chen, Jinyuan
Published: (2025)
NAVIS: Concurrent Search and Update with Low Position-Seeking Overhead in On-SSD Graph-Based Vector Search
by: Song, Jaeyong, et al.
Published: (2026)
by: Song, Jaeyong, et al.
Published: (2026)
Wait-free Replicated Data Types and Fair Reconciliation
by: Kuznetsov, Petr, et al.
Published: (2025)
by: Kuznetsov, Petr, et al.
Published: (2025)
Resilience through Automated Adaptive Configuration for Distribution and Replication
by: Stoller, Scott D., et al.
Published: (2025)
by: Stoller, Scott D., et al.
Published: (2025)
Vertical Atomic Broadcast and Passive Replication (Extended Version)
by: Bravo, Manuel, et al.
Published: (2024)
by: Bravo, Manuel, et al.
Published: (2024)
SWARM: Replicating Shared Disaggregated-Memory Data in No Time
by: Murat, Antoine, et al.
Published: (2024)
by: Murat, Antoine, et al.
Published: (2024)
Orbax: Distributed Checkpointing with JAX
by: Gaffney, Colin, et al.
Published: (2026)
by: Gaffney, Colin, et al.
Published: (2026)
Multi-Resolution Model Fusion for Accelerating the Convolutional Neural Network Training
by: Wang, Kewei, et al.
Published: (2025)
by: Wang, Kewei, et al.
Published: (2025)
Similar Items
-
KaMPIng: Flexible and (Near) Zero-Overhead C++ Bindings for MPI
by: Uhl, Tim Niklas, et al.
Published: (2024) -
Near-Zero-Overhead Freshness for Recommendation Systems via Inference-Side Model Updates
by: Yu, Wenjun, et al.
Published: (2025) -
Optimizing Frequent Checkpointing via Low-Cost Differential for Distributed Training Systems
by: Yao, Chenxuan, et al.
Published: (2025) -
Scaling LLM Inference Beyond Amdahl`s Limits via Eliminating Non-Scalable Overheads
by: Zhao, Alan, et al.
Published: (2026) -
MLTCP: Congestion Control for DNN Training
by: Rajasekaran, Sudarsanan, et al.
Published: (2024)