SPARe: Stacked Parallelism with Adaptive Reordering for Fault-Tolerant LLM Pretraining Systems with 100k+ GPUs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Jin, Chen, Zhonghao, He, Xuhang, Underwood, Robert, Nicolae, Bogdan, Cappello, Franck, Lu, Xiaoyi, Di, Sheng, Zhang, Zheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TurboFFT: Co-Designed High-Performance and Fault-Tolerant Fast Fourier Transform on GPUs
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
DataStates-LLM: Lazy Asynchronous Checkpointing for Large Language Models
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload
von: Liu, Ziyue, et al.
Veröffentlicht: (2026)
von: Liu, Ziyue, et al.
Veröffentlicht: (2026)
DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers
von: Maurya, Avinash, et al.
Veröffentlicht: (2026)
von: Maurya, Avinash, et al.
Veröffentlicht: (2026)
Breaking the Memory Wall: A Study of I/O Patterns and GPU Memory Utilization for Hybrid CPU-GPU Offloaded Optimizers
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
TurboFFT: A High-Performance Fast Fourier Transform with Fault Tolerance on GPU
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
To Compress or Not To Compress: Energy Trade-Offs and Benefits of Lossy Compressed I/O
von: Wilkins, Grant, et al.
Veröffentlicht: (2024)
von: Wilkins, Grant, et al.
Veröffentlicht: (2024)
Training LLMs with Fault Tolerant HSDP on 100,000 GPUs
von: Salpekar, Omkar, et al.
Veröffentlicht: (2026)
von: Salpekar, Omkar, et al.
Veröffentlicht: (2026)
pMSz: A Distributed Parallel Algorithm for Correcting Extrema and Morse Smale Segmentations in Lossy Compression
von: Li, Yuxiao, et al.
Veröffentlicht: (2026)
von: Li, Yuxiao, et al.
Veröffentlicht: (2026)
cuSZ-$i$: High-Ratio Scientific Lossy Compression on GPUs with Optimized Multi-Level Interpolation
von: Liu, Jinyang, et al.
Veröffentlicht: (2023)
von: Liu, Jinyang, et al.
Veröffentlicht: (2023)
FedSZ: Leveraging Error-Bounded Lossy Compression for Federated Learning Communications
von: Wilkins, Grant, et al.
Veröffentlicht: (2023)
von: Wilkins, Grant, et al.
Veröffentlicht: (2023)
DGRO: Diameter-Guided Ring Optimization for Integrated Research Infrastructure Membership
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
FT K-means: A High-Performance K-means on GPU with Fault Tolerance
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
von: Wu, Shixun, et al.
Veröffentlicht: (2024)
Optimizing View Change for Byzantine Fault Tolerance in Parallel Consensus
von: Xie, Yifei, et al.
Veröffentlicht: (2026)
von: Xie, Yifei, et al.
Veröffentlicht: (2026)
Rate-Distortion Bounds for Heterogeneous Random Fields on Finite Lattices
von: Sinha, Sujata, et al.
Veröffentlicht: (2026)
von: Sinha, Sujata, et al.
Veröffentlicht: (2026)
HoSZp: An Efficient Homomorphic Error-bounded Lossy Compressor for Scientific Data
von: Agarwal, Tripti, et al.
Veröffentlicht: (2024)
von: Agarwal, Tripti, et al.
Veröffentlicht: (2024)
Efficient Data-Parallel Continual Learning with Asynchronous Distributed Rehearsal Buffers
von: Bouvier, Thomas, et al.
Veröffentlicht: (2024)
von: Bouvier, Thomas, et al.
Veröffentlicht: (2024)
Parallelizing Maximal Clique Enumeration on GPUs
von: Almasri, Mohammad, et al.
Veröffentlicht: (2022)
von: Almasri, Mohammad, et al.
Veröffentlicht: (2022)
Understanding LLM Checkpoint/Restore I/O Strategies and Patterns
von: Gossman, Mikaila J., et al.
Veröffentlicht: (2025)
von: Gossman, Mikaila J., et al.
Veröffentlicht: (2025)
BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models
von: Wang, Zhengyang, et al.
Veröffentlicht: (2025)
von: Wang, Zhengyang, et al.
Veröffentlicht: (2025)
APEX: Asynchronous Parallel CPU-GPU Execution for Online LLM Inference on Constrained GPUs
von: Fan, Jiakun, et al.
Veröffentlicht: (2025)
von: Fan, Jiakun, et al.
Veröffentlicht: (2025)
Straggler Tolerant and Resilient DL Training on Homogeneous GPUs
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
DeepCQ: General-Purpose Deep-Surrogate Framework for Lossy Compression Quality Prediction
von: Mumenin, Khondoker Mirazul, et al.
Veröffentlicht: (2025)
von: Mumenin, Khondoker Mirazul, et al.
Veröffentlicht: (2025)
Preserving Clusters in Error-Bounded Lossy Compression of Particle Data
von: Ren, Congrong, et al.
Veröffentlicht: (2026)
von: Ren, Congrong, et al.
Veröffentlicht: (2026)
Fault-Tolerant Decentralized Distributed Asynchronous Federated Learning with Adaptive Termination Detection
von: Akkinepally, Phani Sahasra, et al.
Veröffentlicht: (2025)
von: Akkinepally, Phani Sahasra, et al.
Veröffentlicht: (2025)
Beyond Optimal Fault Tolerance
von: Lewis-Pye, Andrew, et al.
Veröffentlicht: (2025)
von: Lewis-Pye, Andrew, et al.
Veröffentlicht: (2025)
S-HPLB: Efficient LLM Attention Serving via Sparsity-Aware Head Parallelism Load Balance
von: Liu, Di, et al.
Veröffentlicht: (2026)
von: Liu, Di, et al.
Veröffentlicht: (2026)
Mitigating Artifacts in Pre-quantization Based Scientific Data Compressors with Quantization-aware Interpolation
von: Jiao, Pu, et al.
Veröffentlicht: (2026)
von: Jiao, Pu, et al.
Veröffentlicht: (2026)
TopoSZp: Lightweight Topology-Aware Error-controlled Compression for Scientific Data
von: Agarwal, Tripti, et al.
Veröffentlicht: (2026)
von: Agarwal, Tripti, et al.
Veröffentlicht: (2026)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
von: Wang, Yuxin, et al.
Veröffentlicht: (2023)
von: Wang, Yuxin, et al.
Veröffentlicht: (2023)
FFCz: Fast Fourier Correction for Spectrum-Preserving Lossy Compression of Scientific Data
von: Ren, Congrong, et al.
Veröffentlicht: (2026)
von: Ren, Congrong, et al.
Veröffentlicht: (2026)
An Adaptive Distributed Stencil Abstraction for GPUs
von: Bhosale, Aditya, et al.
Veröffentlicht: (2025)
von: Bhosale, Aditya, et al.
Veröffentlicht: (2025)
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection
von: Zhou, Yuhang, et al.
Veröffentlicht: (2025)
von: Zhou, Yuhang, et al.
Veröffentlicht: (2025)
Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs
von: Cui, Shengkun, et al.
Veröffentlicht: (2025)
von: Cui, Shengkun, et al.
Veröffentlicht: (2025)
Byzantine Fault Tolerant Causal Ordering
von: Misra, Anshuman, et al.
Veröffentlicht: (2021)
von: Misra, Anshuman, et al.
Veröffentlicht: (2021)
MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant Optimization
von: Hu, Rizhen, et al.
Veröffentlicht: (2025)
von: Hu, Rizhen, et al.
Veröffentlicht: (2025)
PARS3: Parallel Sparse Skew-Symmetric Matrix-Vector Multiplication with Reverse Cuthill-McKee Reordering
von: Yildirim, Selin, et al.
Veröffentlicht: (2024)
von: Yildirim, Selin, et al.
Veröffentlicht: (2024)
LuWu: An End-to-End In-Network Out-of-Core Optimizer for 100B-Scale Model-in-Network Data-Parallel Training on Distributed GPUs
von: Sun, Mo, et al.
Veröffentlicht: (2024)
von: Sun, Mo, et al.
Veröffentlicht: (2024)
IPComp: Interpolation Based Progressive Lossy Compression for Scientific Applications
von: Yang, Zhuoxun, et al.
Veröffentlicht: (2025)
von: Yang, Zhuoxun, et al.
Veröffentlicht: (2025)
Byzantine Fault-Tolerant Min-Max Optimization
von: Liu, Shuo, et al.
Veröffentlicht: (2022)
von: Liu, Shuo, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
TurboFFT: Co-Designed High-Performance and Fault-Tolerant Fast Fourier Transform on GPUs
von: Wu, Shixun, et al.
Veröffentlicht: (2024) -
DataStates-LLM: Lazy Asynchronous Checkpointing for Large Language Models
von: Maurya, Avinash, et al.
Veröffentlicht: (2024) -
ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload
von: Liu, Ziyue, et al.
Veröffentlicht: (2026) -
DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers
von: Maurya, Avinash, et al.
Veröffentlicht: (2026) -
Breaking the Memory Wall: A Study of I/O Patterns and GPU Memory Utilization for Hybrid CPU-GPU Offloaded Optimizers
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)