Understanding Silent Data Corruption in LLM Training
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Jeffrey, Pei, Hengzhi, Lausen, Leonard, Karypis, George |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Photon: Federated LLM Pre-Training
by: Sani, Lorenzo, et al.
Published: (2024)
by: Sani, Lorenzo, et al.
Published: (2024)
SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM Training
by: Jia, Jinda, et al.
Published: (2024)
by: Jia, Jinda, et al.
Published: (2024)
DynaTrain: Fast Online Parallelism Switching for Elastic LLM Training
by: Wang, Yuanqing, et al.
Published: (2026)
by: Wang, Yuanqing, et al.
Published: (2026)
Understanding Stragglers in Large Model Training Using What-if Analysis
by: Lin, Jinkun, et al.
Published: (2025)
by: Lin, Jinkun, et al.
Published: (2025)
HetCCL: Accelerating LLM Training with Heterogeneous GPUs
by: Kim, Heehoon, et al.
Published: (2026)
by: Kim, Heehoon, et al.
Published: (2026)
Unicron: Economizing Self-Healing LLM Training at Scale
by: He, Tao, et al.
Published: (2023)
by: He, Tao, et al.
Published: (2023)
Understanding and Improving Communication Performance in Multi-node LLM Inference
by: Singhania, Prajwal, et al.
Published: (2025)
by: Singhania, Prajwal, et al.
Published: (2025)
GraphStorm: all-in-one graph machine learning framework for industry applications
by: Zheng, Da, et al.
Published: (2024)
by: Zheng, Da, et al.
Published: (2024)
Runtime-Orchestrated Second-Order Optimization for Scalable LLM Training
by: Lu, Yishun, et al.
Published: (2026)
by: Lu, Yishun, et al.
Published: (2026)
Adacc: An Adaptive Framework Unifying Compression and Activation Recomputation for LLM Training
by: Chen, Ping, et al.
Published: (2025)
by: Chen, Ping, et al.
Published: (2025)
Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training
by: Deng, Yangtao, et al.
Published: (2025)
by: Deng, Yangtao, et al.
Published: (2025)
SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
by: Lian, Xinyu, et al.
Published: (2025)
by: Lian, Xinyu, et al.
Published: (2025)
DASH: Deterministic Attention Scheduling for High-throughput Reproducible LLM Training
by: Qiang, Xinwei, et al.
Published: (2026)
by: Qiang, Xinwei, et al.
Published: (2026)
10Cache: Heterogeneous Resource-Aware Tensor Caching and Migration for LLM Training
by: Afroz, Sabiha, et al.
Published: (2025)
by: Afroz, Sabiha, et al.
Published: (2025)
MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training
by: Zhao, Pinxue, et al.
Published: (2024)
by: Zhao, Pinxue, et al.
Published: (2024)
PipeFill: Using GPUs During Bubbles in Pipeline-parallel LLM Training
by: Arfeen, Daiyaan, et al.
Published: (2024)
by: Arfeen, Daiyaan, et al.
Published: (2024)
Empowering Distributed Training with Sparsity-driven Data Synchronization
by: Wang, Zhuang, et al.
Published: (2023)
by: Wang, Zhuang, et al.
Published: (2023)
AsyncHZP: Hierarchical ZeRO Parallelism with Asynchronous Scheduling for Scalable LLM Training
by: Bai, Huawei, et al.
Published: (2025)
by: Bai, Huawei, et al.
Published: (2025)
Nonuniform-Tensor-Parallelism: Mitigating GPU failure impact for Scaled-up LLM Training
by: Arfeen, Daiyaan, et al.
Published: (2025)
by: Arfeen, Daiyaan, et al.
Published: (2025)
A Tabular Schedule Abstraction for Communication-Aware Evaluation of Pipeline-Parallel LLM Training
by: Barley, Daniel, et al.
Published: (2026)
by: Barley, Daniel, et al.
Published: (2026)
Modalities, a PyTorch-native Framework For Large-scale LLM Training and Research
by: Lübbering, Max, et al.
Published: (2026)
by: Lübbering, Max, et al.
Published: (2026)
TensorSocket: Shared Data Loading for Deep Learning Training
by: Robroek, Ties, et al.
Published: (2024)
by: Robroek, Ties, et al.
Published: (2024)
BLoad: Enhancing Neural Network Training with Efficient Sequential Data Handling
by: Ruschel, Raphael, et al.
Published: (2023)
by: Ruschel, Raphael, et al.
Published: (2023)
Comprehensive Evaluation of GNN Training Systems: A Data Management Perspective
by: Yuan, Hao, et al.
Published: (2023)
by: Yuan, Hao, et al.
Published: (2023)
Decoupled Vertical Federated Learning for Practical Training on Vertically Partitioned Data
by: Amalanshu, Avi, et al.
Published: (2024)
by: Amalanshu, Avi, et al.
Published: (2024)
Covenant-72B: Pre-Training a 72B LLM with Trustless Peers Over-the-Internet
by: Lidin, Joel, et al.
Published: (2026)
by: Lidin, Joel, et al.
Published: (2026)
MinatoLoader: Accelerating Machine Learning Training Through Efficient Data Preprocessing
by: Nouaji, Rahma, et al.
Published: (2025)
by: Nouaji, Rahma, et al.
Published: (2025)
DataStates-LLM: Lazy Asynchronous Checkpointing for Large Language Models
by: Maurya, Avinash, et al.
Published: (2024)
by: Maurya, Avinash, et al.
Published: (2024)
FedStaleWeight: Buffered Asynchronous Federated Learning with Fair Aggregation via Staleness Reweighting
by: Ma, Jeffrey, et al.
Published: (2024)
by: Ma, Jeffrey, et al.
Published: (2024)
BatchWeave: A Consistent Object-Store-Native Data Plane for Large Foundation Model Training
by: Sun, Ting, et al.
Published: (2026)
by: Sun, Ting, et al.
Published: (2026)
Hydraulis: Balancing Large Transformer Model Training via Co-designing Parallel Strategies and Data Assignment
by: Li, Haoyang, et al.
Published: (2024)
by: Li, Haoyang, et al.
Published: (2024)
LSM-GNN: Large-scale Storage-based Multi-GPU GNN Training by Optimizing Data Transfer Scheme
by: Park, Jeongmin Brian, et al.
Published: (2024)
by: Park, Jeongmin Brian, et al.
Published: (2024)
Distributed Training under Packet Loss
by: Weintraub, Erez, et al.
Published: (2025)
by: Weintraub, Erez, et al.
Published: (2025)
Echo: Simulating Distributed Training At Scale
by: Feng, Yicheng, et al.
Published: (2024)
by: Feng, Yicheng, et al.
Published: (2024)
ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving
by: Qiao, Yifan, et al.
Published: (2024)
by: Qiao, Yifan, et al.
Published: (2024)
ScaleLLM: A Resource-Frugal LLM Serving Framework by Optimizing End-to-End Efficiency
by: Yao, Yuhang, et al.
Published: (2024)
by: Yao, Yuhang, et al.
Published: (2024)
FedAST: Federated Asynchronous Simultaneous Training
by: Askin, Baris, et al.
Published: (2024)
by: Askin, Baris, et al.
Published: (2024)
Reducing Energy Bloat in Large Model Training
by: Chung, Jae-Won, et al.
Published: (2023)
by: Chung, Jae-Won, et al.
Published: (2023)
MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production
by: Jin, Chao, et al.
Published: (2025)
by: Jin, Chao, et al.
Published: (2025)
On Harnessing Idle Compute at the Edge for Foundation Model Training
by: Xue, Leyang, et al.
Published: (2025)
by: Xue, Leyang, et al.
Published: (2025)
Similar Items
-
Photon: Federated LLM Pre-Training
by: Sani, Lorenzo, et al.
Published: (2024) -
SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM Training
by: Jia, Jinda, et al.
Published: (2024) -
DynaTrain: Fast Online Parallelism Switching for Elastic LLM Training
by: Wang, Yuanqing, et al.
Published: (2026) -
Understanding Stragglers in Large Model Training Using What-if Analysis
by: Lin, Jinkun, et al.
Published: (2025) -
HetCCL: Accelerating LLM Training with Heterogeneous GPUs
by: Kim, Heehoon, et al.
Published: (2026)