ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Ziyue, Wang, Zhengyang, Zhang, Ruijie, Maurya, Avinash, Zhou, Hui, Hovland, Paul, Di, Sheng, Cappello, Franck, Nicolae, Bogdan, Zhang, Zheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models
von: Wang, Zhengyang, et al.
Veröffentlicht: (2025)
von: Wang, Zhengyang, et al.
Veröffentlicht: (2025)
DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers
von: Maurya, Avinash, et al.
Veröffentlicht: (2026)
von: Maurya, Avinash, et al.
Veröffentlicht: (2026)
Breaking the Memory Wall: A Study of I/O Patterns and GPU Memory Utilization for Hybrid CPU-GPU Offloaded Optimizers
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
DataStates-LLM: Lazy Asynchronous Checkpointing for Large Language Models
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
MLP-Offload: Multi-Level, Multi-Path Offloading for LLM Pre-training to Break the GPU Memory Wall
von: Maurya, Avinash, et al.
Veröffentlicht: (2025)
von: Maurya, Avinash, et al.
Veröffentlicht: (2025)
SPARe: Stacked Parallelism with Adaptive Reordering for Fault-Tolerant LLM Pretraining Systems with 100k+ GPUs
von: Lee, Jin, et al.
Veröffentlicht: (2026)
von: Lee, Jin, et al.
Veröffentlicht: (2026)
Deep Optimizer States: Towards Scalable Training of Transformer Models Using Interleaved Offloading
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles
von: Arif, Moiz, et al.
Veröffentlicht: (2026)
von: Arif, Moiz, et al.
Veröffentlicht: (2026)
Understanding LLM Checkpoint/Restore I/O Strategies and Patterns
von: Gossman, Mikaila J., et al.
Veröffentlicht: (2025)
von: Gossman, Mikaila J., et al.
Veröffentlicht: (2025)
TurboFFT: Co-Designed High-Performance and Fault-Tolerant Fast Fourier Transform on GPUs
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
TurboFFT: A High-Performance Fast Fourier Transform with Fault Tolerance on GPU
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
Mitigating Artifacts in Pre-quantization Based Scientific Data Compressors with Quantization-aware Interpolation
von: Jiao, Pu, et al.
Veröffentlicht: (2026)
von: Jiao, Pu, et al.
Veröffentlicht: (2026)
FT K-means: A High-Performance K-means on GPU with Fault Tolerance
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
DGRO: Diameter-Guided Ring Optimization for Integrated Research Infrastructure Membership
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
To Compress or Not To Compress: Energy Trade-Offs and Benefits of Lossy Compressed I/O
von: Wilkins, Grant, et al.
Veröffentlicht: (2024)
von: Wilkins, Grant, et al.
Veröffentlicht: (2024)
Straggler Tolerant and Resilient DL Training on Homogeneous GPUs
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
On Fault Tolerance of Data Storage Systems: A Holistic Perspective
von: Zheng, Mai, et al.
Veröffentlicht: (2025)
von: Zheng, Mai, et al.
Veröffentlicht: (2025)
QPET: A Versatile and Portable Quantity-of-Interest-Preservation Framework for Error-Bounded Lossy Compression
von: Liu, Jinyang, et al.
Veröffentlicht: (2024)
von: Liu, Jinyang, et al.
Veröffentlicht: (2024)
Beyond Optimal Fault Tolerance
von: Lewis-Pye, Andrew, et al.
Veröffentlicht: (2025)
von: Lewis-Pye, Andrew, et al.
Veröffentlicht: (2025)
Byzantine Fault Tolerant Causal Ordering
von: Misra, Anshuman, et al.
Veröffentlicht: (2021)
von: Misra, Anshuman, et al.
Veröffentlicht: (2021)
An Optimized Error-controlled MPI Collective Framework Integrated with Lossy Compression
von: Huang, Jiajun, et al.
Veröffentlicht: (2023)
von: Huang, Jiajun, et al.
Veröffentlicht: (2023)
ZCCL: Significantly Improving Collective Communication With Error-Bounded Lossy Compression
von: Huang, Jiajun, et al.
Veröffentlicht: (2025)
von: Huang, Jiajun, et al.
Veröffentlicht: (2025)
Preserving Clusters in Error-Bounded Lossy Compression of Particle Data
von: Ren, Congrong, et al.
Veröffentlicht: (2026)
von: Ren, Congrong, et al.
Veröffentlicht: (2026)
Byzantine Fault-Tolerant Min-Max Optimization
von: Liu, Shuo, et al.
Veröffentlicht: (2022)
von: Liu, Shuo, et al.
Veröffentlicht: (2022)
Approximate Byzantine Fault-Tolerance in Distributed Optimization
von: Liu, Shuo, et al.
Veröffentlicht: (2021)
von: Liu, Shuo, et al.
Veröffentlicht: (2021)
Optimal Fault-Tolerant Dispersion on Oriented Grids
von: Banerjee, Rik, et al.
Veröffentlicht: (2024)
von: Banerjee, Rik, et al.
Veröffentlicht: (2024)
Probabilistic Byzantine Fault Tolerance (Extended Version)
von: Avelãs, Diogo, et al.
Veröffentlicht: (2024)
von: Avelãs, Diogo, et al.
Veröffentlicht: (2024)
Stabl: Blockchain Fault Tolerance
von: Gramoli, Vincent, et al.
Veröffentlicht: (2024)
von: Gramoli, Vincent, et al.
Veröffentlicht: (2024)
IPComp: Interpolation Based Progressive Lossy Compression for Scientific Applications
von: Yang, Zhuoxun, et al.
Veröffentlicht: (2025)
von: Yang, Zhuoxun, et al.
Veröffentlicht: (2025)
Wilkins: HPC In Situ Workflows Made Easy
von: Yildiz, Orcun, et al.
Veröffentlicht: (2024)
von: Yildiz, Orcun, et al.
Veröffentlicht: (2024)
gZCCL: Compression-Accelerated Collective Communication Framework for GPU Clusters
von: Huang, Jiajun, et al.
Veröffentlicht: (2023)
von: Huang, Jiajun, et al.
Veröffentlicht: (2023)
Optimizing Robot Dispersion on Grids: with and without Fault Tolerance
von: Banerjee, Rik, et al.
Veröffentlicht: (2024)
von: Banerjee, Rik, et al.
Veröffentlicht: (2024)
VBFT: Veloce Byzantine Fault Tolerant Consensus for Blockchains
von: Jalalzai, Mohammad M., et al.
Veröffentlicht: (2023)
von: Jalalzai, Mohammad M., et al.
Veröffentlicht: (2023)
A Fault Tolerance Mechanism for Hybrid Scientific Workflows
von: Mulone, Alberto, et al.
Veröffentlicht: (2024)
von: Mulone, Alberto, et al.
Veröffentlicht: (2024)
Asynchronous Fault-Tolerant Distributed Proper Coloring of Graphs
von: Balliu, Alkida, et al.
Veröffentlicht: (2024)
von: Balliu, Alkida, et al.
Veröffentlicht: (2024)
Arma: Byzantine Fault Tolerant Consensus with Horizontal Scalability
von: Manevich, Yacov, et al.
Veröffentlicht: (2024)
von: Manevich, Yacov, et al.
Veröffentlicht: (2024)
The Case for ABI Interoperability in a Fault Tolerant MPI
von: Xu, Yao, et al.
Veröffentlicht: (2025)
von: Xu, Yao, et al.
Veröffentlicht: (2025)
Half a Century of Distributed Byzantine Fault-Tolerant Consensus: Design Principles and Evolutionary Pathways
von: Wu, Huanyu, et al.
Veröffentlicht: (2024)
von: Wu, Huanyu, et al.
Veröffentlicht: (2024)
TopoSZp: Lightweight Topology-Aware Error-controlled Compression for Scientific Data
von: Agarwal, Tripti, et al.
Veröffentlicht: (2026)
von: Agarwal, Tripti, et al.
Veröffentlicht: (2026)
A Byzantine Fault Tolerance Approach towards AI Safety
von: deVadoss, John, et al.
Veröffentlicht: (2025)
von: deVadoss, John, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models
von: Wang, Zhengyang, et al.
Veröffentlicht: (2025) -
DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers
von: Maurya, Avinash, et al.
Veröffentlicht: (2026) -
Breaking the Memory Wall: A Study of I/O Patterns and GPU Memory Utilization for Hybrid CPU-GPU Offloaded Optimizers
von: Maurya, Avinash, et al.
Veröffentlicht: (2024) -
DataStates-LLM: Lazy Asynchronous Checkpointing for Large Language Models
von: Maurya, Avinash, et al.
Veröffentlicht: (2024) -
MLP-Offload: Multi-Level, Multi-Path Offloading for LLM Pre-training to Break the GPU Memory Wall
von: Maurya, Avinash, et al.
Veröffentlicht: (2025)