Understanding LLM Checkpoint/Restore I/O Strategies and Patterns
Fuente:
arXiv
Saved in:
| Main Authors: | Gossman, Mikaila J., Maurya, Avinash, Nicolae, Bogdan, Calhoun, Jon C. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers
by: Maurya, Avinash, et al.
Published: (2026)
by: Maurya, Avinash, et al.
Published: (2026)
DataStates-LLM: Lazy Asynchronous Checkpointing for Large Language Models
by: Maurya, Avinash, et al.
Published: (2024)
by: Maurya, Avinash, et al.
Published: (2024)
Breaking the Memory Wall: A Study of I/O Patterns and GPU Memory Utilization for Hybrid CPU-GPU Offloaded Optimizers
by: Maurya, Avinash, et al.
Published: (2024)
by: Maurya, Avinash, et al.
Published: (2024)
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles
by: Arif, Moiz, et al.
Published: (2026)
by: Arif, Moiz, et al.
Published: (2026)
To Compress or Not To Compress: Energy Trade-Offs and Benefits of Lossy Compressed I/O
by: Wilkins, Grant, et al.
Published: (2024)
by: Wilkins, Grant, et al.
Published: (2024)
ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload
by: Liu, Ziyue, et al.
Published: (2026)
by: Liu, Ziyue, et al.
Published: (2026)
BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models
by: Wang, Zhengyang, et al.
Published: (2025)
by: Wang, Zhengyang, et al.
Published: (2025)
CRIU -- Checkpoint Restore in Userspace for computational simulations and scientific applications
by: Andrijauskas, Fabio, et al.
Published: (2024)
by: Andrijauskas, Fabio, et al.
Published: (2024)
Wilkins: HPC In Situ Workflows Made Easy
by: Yildiz, Orcun, et al.
Published: (2024)
by: Yildiz, Orcun, et al.
Published: (2024)
Efficient LLM Inference with Activation Checkpointing and Hybrid Caching
by: Lee, Sanghyeon, et al.
Published: (2025)
by: Lee, Sanghyeon, et al.
Published: (2025)
Efficient Data-Parallel Continual Learning with Asynchronous Distributed Rehearsal Buffers
by: Bouvier, Thomas, et al.
Published: (2024)
by: Bouvier, Thomas, et al.
Published: (2024)
Deep Optimizer States: Towards Scalable Training of Transformer Models Using Interleaved Offloading
by: Maurya, Avinash, et al.
Published: (2024)
by: Maurya, Avinash, et al.
Published: (2024)
Architectural Foundations for Checkpointing and Restoration in Quantum HPC Systems
by: Guan, Qiang, et al.
Published: (2026)
by: Guan, Qiang, et al.
Published: (2026)
Kavier: Exploring Performance, Sustainability, and Efficiency of LLM Ecosystems under Inference through Cache-Aware Discrete-Event Simulation
by: Nicolae, Radu, et al.
Published: (2026)
by: Nicolae, Radu, et al.
Published: (2026)
MLP-Offload: Multi-Level, Multi-Path Offloading for LLM Pre-training to Break the GPU Memory Wall
by: Maurya, Avinash, et al.
Published: (2025)
by: Maurya, Avinash, et al.
Published: (2025)
Fast State Restoration in LLM Serving with HCache
by: Gao, Shiwei, et al.
Published: (2024)
by: Gao, Shiwei, et al.
Published: (2024)
Asynchronous Checkpoint for Eventually Consistent Databases
by: Ravishankar, Raaghav, et al.
Published: (2025)
by: Ravishankar, Raaghav, et al.
Published: (2025)
FedSZ: Leveraging Error-Bounded Lossy Compression for Federated Learning Communications
by: Wilkins, Grant, et al.
Published: (2023)
by: Wilkins, Grant, et al.
Published: (2023)
CRIUgpu: Transparent Checkpointing of GPU-Accelerated Workloads
by: Stoyanov, Radostin, et al.
Published: (2025)
by: Stoyanov, Radostin, et al.
Published: (2025)
Scrutinizing Variables for Checkpoint Using Automatic Differentiation
by: Huang, Xin, et al.
Published: (2026)
by: Huang, Xin, et al.
Published: (2026)
Optimal Checkpoint Interval with Availability as an Objective Function
by: Saxena, Nirmal Raj, et al.
Published: (2024)
by: Saxena, Nirmal Raj, et al.
Published: (2024)
Checkpoint and Restart: An Energy Consumption Characterization in Clusters
by: Moran, Marina, et al.
Published: (2024)
by: Moran, Marina, et al.
Published: (2024)
PhoenixOS: Concurrent OS-level GPU Checkpoint and Restore with Validated Speculation
by: Wei, Xingda, et al.
Published: (2024)
by: Wei, Xingda, et al.
Published: (2024)
Sparse Checkpointing for Fast and Reliable MoE Training
by: Gandhi, Swapnil, et al.
Published: (2024)
by: Gandhi, Swapnil, et al.
Published: (2024)
All is Not Lost: LLM Recovery without Checkpoints
by: Blagoev, Nikolay, et al.
Published: (2025)
by: Blagoev, Nikolay, et al.
Published: (2025)
Checkmate: Zero-Overhead Model Checkpointing via Network Gradient Replication
by: Bhardwaj, Ankit, et al.
Published: (2025)
by: Bhardwaj, Ankit, et al.
Published: (2025)
CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration
by: Nian, Sean, et al.
Published: (2026)
by: Nian, Sean, et al.
Published: (2026)
M3SA: Exploring Datacenter Performance and Climate-Impact with Multi- and Meta-Model Simulation and Analysis
by: Nicolae, Radu, et al.
Published: (2026)
by: Nicolae, Radu, et al.
Published: (2026)
Optimizing Frequent Checkpointing via Low-Cost Differential for Distributed Training Systems
by: Yao, Chenxuan, et al.
Published: (2025)
by: Yao, Chenxuan, et al.
Published: (2025)
SPARe: Stacked Parallelism with Adaptive Reordering for Fault-Tolerant LLM Pretraining Systems with 100k+ GPUs
by: Lee, Jin, et al.
Published: (2026)
by: Lee, Jin, et al.
Published: (2026)
In-Transit Data Transport Strategies for Coupled AI-Simulation Workflow Patterns
by: Tummalapalli, Harikrishna, et al.
Published: (2025)
by: Tummalapalli, Harikrishna, et al.
Published: (2025)
LLMTailor: A Layer-wise Tailoring Tool for Efficient Checkpointing of Large Language Models
by: Sun, Minqiu, et al.
Published: (2026)
by: Sun, Minqiu, et al.
Published: (2026)
D-Rex: Heterogeneity-Aware Reliability Framework and Adaptive Algorithms for Distributed Storage
by: Gonthier, Maxime, et al.
Published: (2025)
by: Gonthier, Maxime, et al.
Published: (2025)
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
by: Arya, Mayank, et al.
Published: (2025)
by: Arya, Mayank, et al.
Published: (2025)
CheckMate: Evaluating Checkpointing Protocols for Streaming Dataflows
by: Siachamis, George, et al.
Published: (2024)
by: Siachamis, George, et al.
Published: (2024)
OpenDT: Exploring Datacenter Performance and Sustainability with a Self-Calibrating Digital Twin
by: Nicolae, Radu, et al.
Published: (2026)
by: Nicolae, Radu, et al.
Published: (2026)
An Overview on the Landscape of Self-Adaptive Cloud Design and Operation Patterns: Goals, Strategies, Tooling, Evaluation, and Dataset Perspectives
by: Angelis, Apostolos, et al.
Published: (2025)
by: Angelis, Apostolos, et al.
Published: (2025)
ParaLog: Consistent Host-side Logging for Parallel Checkpoints
by: Chien, Steven W. D., et al.
Published: (2024)
by: Chien, Steven W. D., et al.
Published: (2024)
Orbax: Distributed Checkpointing with JAX
by: Gaffney, Colin, et al.
Published: (2026)
by: Gaffney, Colin, et al.
Published: (2026)
Tightening I/O Lower Bounds through the Hourglass Dependency Pattern
by: Eyraud-Dubois, Lionel, et al.
Published: (2024)
by: Eyraud-Dubois, Lionel, et al.
Published: (2024)
Similar Items
-
DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers
by: Maurya, Avinash, et al.
Published: (2026) -
DataStates-LLM: Lazy Asynchronous Checkpointing for Large Language Models
by: Maurya, Avinash, et al.
Published: (2024) -
Breaking the Memory Wall: A Study of I/O Patterns and GPU Memory Utilization for Hybrid CPU-GPU Offloaded Optimizers
by: Maurya, Avinash, et al.
Published: (2024) -
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles
by: Arif, Moiz, et al.
Published: (2026) -
To Compress or Not To Compress: Energy Trade-Offs and Benefits of Lossy Compressed I/O
by: Wilkins, Grant, et al.
Published: (2024)