Scrutinizing Variables for Checkpoint Using Automatic Differentiation
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Xin, Zhang, Weiping, Meng, Shiman, Xu, Wubiao, Fu, Xiang, Guo, Luanzheng, Sato, Kento |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Distributed Order Recording Techniques for Efficient Record-and-Replay of Multi-threaded Programs
by: Fu, Xiang, et al.
Published: (2026)
by: Fu, Xiang, et al.
Published: (2026)
LLMTailor: A Layer-wise Tailoring Tool for Efficient Checkpointing of Large Language Models
by: Sun, Minqiu, et al.
Published: (2026)
by: Sun, Minqiu, et al.
Published: (2026)
ParaLog: Consistent Host-side Logging for Parallel Checkpoints
by: Chien, Steven W. D., et al.
Published: (2024)
by: Chien, Steven W. D., et al.
Published: (2024)
AutoCheck: Automatically Identifying Variables for Checkpointing by Data Dependency Analysis
by: Fu, Xiang, et al.
Published: (2024)
by: Fu, Xiang, et al.
Published: (2024)
On The Reproducibility Limitations of RAG Systems
by: Wang, Baiqiang, et al.
Published: (2025)
by: Wang, Baiqiang, et al.
Published: (2025)
QoSFlow: Ensuring Service Quality of Distributed Workflows Using Interpretable Sensitivity Models
by: Rashid, Md Hasanur, et al.
Published: (2026)
by: Rashid, Md Hasanur, et al.
Published: (2026)
PowerTrip: Exploiting Federated Heterogeneous Datacenter Power for Distributed ML Training
by: Mehboob, Talha, et al.
Published: (2025)
by: Mehboob, Talha, et al.
Published: (2025)
Optimizing Frequent Checkpointing via Low-Cost Differential for Distributed Training Systems
by: Yao, Chenxuan, et al.
Published: (2025)
by: Yao, Chenxuan, et al.
Published: (2025)
Understanding Power Consumption Metric on Heterogeneous Memory Systems
by: Proaño, Andrès Rubio, et al.
Published: (2024)
by: Proaño, Andrès Rubio, et al.
Published: (2024)
Asynchronous Checkpoint for Eventually Consistent Databases
by: Ravishankar, Raaghav, et al.
Published: (2025)
by: Ravishankar, Raaghav, et al.
Published: (2025)
Leader Rotation Is Not Enough: Scrutinizing Leadership Democracy of Chained BFT Consensus
by: Tang, Yining, et al.
Published: (2025)
by: Tang, Yining, et al.
Published: (2025)
CRIUgpu: Transparent Checkpointing of GPU-Accelerated Workloads
by: Stoyanov, Radostin, et al.
Published: (2025)
by: Stoyanov, Radostin, et al.
Published: (2025)
Optimal Checkpoint Interval with Availability as an Objective Function
by: Saxena, Nirmal Raj, et al.
Published: (2024)
by: Saxena, Nirmal Raj, et al.
Published: (2024)
Checkpoint and Restart: An Energy Consumption Characterization in Clusters
by: Moran, Marina, et al.
Published: (2024)
by: Moran, Marina, et al.
Published: (2024)
Sparse Checkpointing for Fast and Reliable MoE Training
by: Gandhi, Swapnil, et al.
Published: (2024)
by: Gandhi, Swapnil, et al.
Published: (2024)
Efficient LLM Inference with Activation Checkpointing and Hybrid Caching
by: Lee, Sanghyeon, et al.
Published: (2025)
by: Lee, Sanghyeon, et al.
Published: (2025)
CRIU -- Checkpoint Restore in Userspace for computational simulations and scientific applications
by: Andrijauskas, Fabio, et al.
Published: (2024)
by: Andrijauskas, Fabio, et al.
Published: (2024)
Understanding LLM Checkpoint/Restore I/O Strategies and Patterns
by: Gossman, Mikaila J., et al.
Published: (2025)
by: Gossman, Mikaila J., et al.
Published: (2025)
ZEUS: An Efficient GPU Optimization Method Integrating PSO, BFGS, and Automatic Differentiation
by: Soos, Dominik, et al.
Published: (2026)
by: Soos, Dominik, et al.
Published: (2026)
Checkmate: Zero-Overhead Model Checkpointing via Network Gradient Replication
by: Bhardwaj, Ankit, et al.
Published: (2025)
by: Bhardwaj, Ankit, et al.
Published: (2025)
DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers
by: Maurya, Avinash, et al.
Published: (2026)
by: Maurya, Avinash, et al.
Published: (2026)
ADELIA: Automatic Differentiation for Efficient Laplace Inference Approximations
by: Boudaoud, Afif, et al.
Published: (2026)
by: Boudaoud, Afif, et al.
Published: (2026)
Union: An Automatic Workload Manager for Accelerating Network Simulation
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
by: Wang, Yuxin, et al.
Published: (2023)
by: Wang, Yuxin, et al.
Published: (2023)
CheckMate: Evaluating Checkpointing Protocols for Streaming Dataflows
by: Siachamis, George, et al.
Published: (2024)
by: Siachamis, George, et al.
Published: (2024)
Universal Checkpointing: A Flexible and Efficient Distributed Checkpointing System for Large-Scale DNN Training with Reconfigurable Parallelis
by: Lian, Xinyu, et al.
Published: (2024)
by: Lian, Xinyu, et al.
Published: (2024)
An Integrated (Crop Model, Cloud and Big Data Analytic) Framework to support Agriculture Activity Monitoring System
by: Akhter, Shamim, et al.
Published: (2024)
by: Akhter, Shamim, et al.
Published: (2024)
Orbax: Distributed Checkpointing with JAX
by: Gaffney, Colin, et al.
Published: (2026)
by: Gaffney, Colin, et al.
Published: (2026)
TierCheck: Tiered Checkpointing for Fault Tolerance in Large Language Model Training
by: Han, Shujie, et al.
Published: (2026)
by: Han, Shujie, et al.
Published: (2026)
Towards Fully Automatic Distributed Lower Bounds
by: Balliu, Alkida, et al.
Published: (2024)
by: Balliu, Alkida, et al.
Published: (2024)
Automatic Tracing in Task-Based Runtime Systems
by: Yadav, Rohan, et al.
Published: (2024)
by: Yadav, Rohan, et al.
Published: (2024)
Automatic Metadata Capture and Processing for High-Performance Workflows
by: Shpilker, Polina, et al.
Published: (2025)
by: Shpilker, Polina, et al.
Published: (2025)
Galvatron: Automatic Distributed Training for Large Transformer Models
by: Gumaan, Esmail
Published: (2025)
by: Gumaan, Esmail
Published: (2025)
Kilometer-Level Coupled Modeling Using 40 Million Cores: An Eight-Year Journey of Model Development
by: Duan, Xiaohui, et al.
Published: (2024)
by: Duan, Xiaohui, et al.
Published: (2024)
All is Not Lost: LLM Recovery without Checkpoints
by: Blagoev, Nikolay, et al.
Published: (2025)
by: Blagoev, Nikolay, et al.
Published: (2025)
Warp-STAR: High-performance, Differentiable GPU-Accelerated Static Timing Analysis through Warp-oriented Parallel Orchestration
by: Huang, En-Ming, et al.
Published: (2026)
by: Huang, En-Ming, et al.
Published: (2026)
Federated Automatic Differentiation
by: Rush, Keith, et al.
Published: (2023)
by: Rush, Keith, et al.
Published: (2023)
Paradigm Shift in Infrastructure Inspection Technology: Leveraging High-performance Imaging and Advanced AI Analytics to Inspect Road Infrastructure
by: Wu, Du, et al.
Published: (2025)
by: Wu, Du, et al.
Published: (2025)
Proposal of Automatic Offloading Method in Mixed Offloading Destination Environment
by: Yamato, Yoji
Published: (2020)
by: Yamato, Yoji
Published: (2020)
PhoenixOS: Concurrent OS-level GPU Checkpoint and Restore with Validated Speculation
by: Wei, Xingda, et al.
Published: (2024)
by: Wei, Xingda, et al.
Published: (2024)
Similar Items
-
Distributed Order Recording Techniques for Efficient Record-and-Replay of Multi-threaded Programs
by: Fu, Xiang, et al.
Published: (2026) -
LLMTailor: A Layer-wise Tailoring Tool for Efficient Checkpointing of Large Language Models
by: Sun, Minqiu, et al.
Published: (2026) -
ParaLog: Consistent Host-side Logging for Parallel Checkpoints
by: Chien, Steven W. D., et al.
Published: (2024) -
AutoCheck: Automatically Identifying Variables for Checkpointing by Data Dependency Analysis
by: Fu, Xiang, et al.
Published: (2024) -
On The Reproducibility Limitations of RAG Systems
by: Wang, Baiqiang, et al.
Published: (2025)