Training Through Failure: Effects of Data Consistency in Parallel Machine Learning Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cao, Ray, Luo, Sherry, Gan, Steve, Jinesh, Sujeeth |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Collaborative Split Federated Learning with Parallel Training and Aggregation
von: Papageorgiou, Yiannis, et al.
Veröffentlicht: (2025)
von: Papageorgiou, Yiannis, et al.
Veröffentlicht: (2025)
Accelerating Long-Tail Generation in Synchronous RLHF Training via Adaptive Tensor Parallelism
von: Zhao, Long, et al.
Veröffentlicht: (2026)
von: Zhao, Long, et al.
Veröffentlicht: (2026)
TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training
von: Liu, Man, et al.
Veröffentlicht: (2026)
von: Liu, Man, et al.
Veröffentlicht: (2026)
FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs
von: Zhang, Haijun, et al.
Veröffentlicht: (2025)
von: Zhang, Haijun, et al.
Veröffentlicht: (2025)
InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training
von: Wang, Shiju, et al.
Veröffentlicht: (2025)
von: Wang, Shiju, et al.
Veröffentlicht: (2025)
Mist: Efficient Distributed Training of Large Language Models via Memory-Parallelism Co-Optimization
von: Zhu, Zhanda, et al.
Veröffentlicht: (2025)
von: Zhu, Zhanda, et al.
Veröffentlicht: (2025)
MegaScale-Data: Scaling Dataloader for Multisource Large Foundation Model Training
von: Zhao, Juntao, et al.
Veröffentlicht: (2025)
von: Zhao, Juntao, et al.
Veröffentlicht: (2025)
Cloudless-Training: A Framework to Improve Efficiency of Geo-Distributed ML Training
von: Tan, Wenting, et al.
Veröffentlicht: (2023)
von: Tan, Wenting, et al.
Veröffentlicht: (2023)
Training Overhead Ratio: A Practical Reliability Metric for Large Language Model Training Systems
von: Lu, Ning, et al.
Veröffentlicht: (2024)
von: Lu, Ning, et al.
Veröffentlicht: (2024)
GraphPipe: Improving Performance and Scalability of DNN Training with Graph Pipeline Parallelism
von: Jeon, Byungsoo, et al.
Veröffentlicht: (2024)
von: Jeon, Byungsoo, et al.
Veröffentlicht: (2024)
Research on Model Parallelism and Data Parallelism Optimization Methods in Large Language Model-Based Recommendation Systems
von: Yang, Haowei, et al.
Veröffentlicht: (2025)
von: Yang, Haowei, et al.
Veröffentlicht: (2025)
Scalable and Adaptive Parallel Training of Graph Transformer on Large Graphs
von: Lin, Jun-Liang, et al.
Veröffentlicht: (2026)
von: Lin, Jun-Liang, et al.
Veröffentlicht: (2026)
ResiHP: Taming LLM Training Failures with Dynamic Hybrid Parallelism
von: Ma, Tenghui, et al.
Veröffentlicht: (2026)
von: Ma, Tenghui, et al.
Veröffentlicht: (2026)
Revisiting Parameter Server in LLM Post-Training
von: Wan, Xinyi, et al.
Veröffentlicht: (2026)
von: Wan, Xinyi, et al.
Veröffentlicht: (2026)
Optimizing Data Distribution and Kernel Performance for Efficient Training of Chemistry Foundation Models: A Case Study with MACE
von: Firoz, Jesun, et al.
Veröffentlicht: (2025)
von: Firoz, Jesun, et al.
Veröffentlicht: (2025)
OrchMLLM: Orchestrate Multimodal Data with Batch Post-Balancing to Accelerate Multimodal Large Language Model Training
von: Zheng, Yijie, et al.
Veröffentlicht: (2025)
von: Zheng, Yijie, et al.
Veröffentlicht: (2025)
OptPipe: Memory- and Scheduling-Optimized Pipeline Parallelism for LLM Training
von: Li, Hongpei, et al.
Veröffentlicht: (2025)
von: Li, Hongpei, et al.
Veröffentlicht: (2025)
A Parallel Alternative for Energy-Efficient Neural Network Training and Inferencing
von: Seal, Sudip K., et al.
Veröffentlicht: (2025)
von: Seal, Sudip K., et al.
Veröffentlicht: (2025)
SimpleFSDP: Simpler Fully Sharded Data Parallel with torch.compile
von: Zhang, Ruisi, et al.
Veröffentlicht: (2024)
von: Zhang, Ruisi, et al.
Veröffentlicht: (2024)
PacTrain: Pruning and Adaptive Sparse Gradient Compression for Efficient Collective Communication in Distributed Deep Learning
von: Wang, Yisu, et al.
Veröffentlicht: (2025)
von: Wang, Yisu, et al.
Veröffentlicht: (2025)
Balanced and Elastic End-to-end Training of Dynamic LLMs
von: Wahib, Mohamed, et al.
Veröffentlicht: (2025)
von: Wahib, Mohamed, et al.
Veröffentlicht: (2025)
Ensemble Method for System Failure Detection Using Large-Scale Telemetry Data
von: Mudgal, Priyanka, et al.
Veröffentlicht: (2024)
von: Mudgal, Priyanka, et al.
Veröffentlicht: (2024)
Analyzing the Impact of Participant Failures in Cross-Silo Federated Learning
von: Stricker, Fabian, et al.
Veröffentlicht: (2025)
von: Stricker, Fabian, et al.
Veröffentlicht: (2025)
Robust Synchronisation for Federated Learning in The Face of Correlated Device Failure
von: Behfar, Stefan, et al.
Veröffentlicht: (2026)
von: Behfar, Stefan, et al.
Veröffentlicht: (2026)
JanusPipe: Efficient Pipeline Parallel Training for Machine Learning Interatomic Potentials
von: Wang, Hongyu, et al.
Veröffentlicht: (2026)
von: Wang, Hongyu, et al.
Veröffentlicht: (2026)
A Few GPUs, A Whole Lotta Scale: Faithful LLM Training Emulation with PrismLLM
von: Xi, Shaoke, et al.
Veröffentlicht: (2026)
von: Xi, Shaoke, et al.
Veröffentlicht: (2026)
BitPipe: Bidirectional Interleaved Pipeline Parallelism for Accelerating Large Models Training
von: Wu, Houming, et al.
Veröffentlicht: (2024)
von: Wu, Houming, et al.
Veröffentlicht: (2024)
Training LLMs with Fault Tolerant HSDP on 100,000 GPUs
von: Salpekar, Omkar, et al.
Veröffentlicht: (2026)
von: Salpekar, Omkar, et al.
Veröffentlicht: (2026)
High-Dimensional Data Processing: Benchmarking Machine Learning and Deep Learning Architectures in Local and Distributed Environments
von: Rodriguez, Julian, et al.
Veröffentlicht: (2025)
von: Rodriguez, Julian, et al.
Veröffentlicht: (2025)
DWDP: Distributed Weight Data Parallelism for High-Performance LLM Inference on NVL72
von: Li, Wanqian, et al.
Veröffentlicht: (2026)
von: Li, Wanqian, et al.
Veröffentlicht: (2026)
Towards Scalable GPU-Accelerated SNN Training via Temporal Fusion
von: Li, Yanchen, et al.
Veröffentlicht: (2024)
von: Li, Yanchen, et al.
Veröffentlicht: (2024)
Accelerating Large Language Model Training with Hybrid GPU-based Compression
von: Xu, Lang, et al.
Veröffentlicht: (2024)
von: Xu, Lang, et al.
Veröffentlicht: (2024)
TensorHub: Scalable and Elastic Weight Transfer for LLM RL Training
von: Ye, Chenhao, et al.
Veröffentlicht: (2026)
von: Ye, Chenhao, et al.
Veröffentlicht: (2026)
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning
von: Xu, Lang, et al.
Veröffentlicht: (2025)
von: Xu, Lang, et al.
Veröffentlicht: (2025)
ModTrans: Translating Real-world Models for Distributed Training Simulator
von: Lyu, Yi
Veröffentlicht: (2026)
von: Lyu, Yi
Veröffentlicht: (2026)
Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence
von: Chen, Xinquan, et al.
Veröffentlicht: (2026)
von: Chen, Xinquan, et al.
Veröffentlicht: (2026)
Lumos: Efficient Performance Modeling and Estimation for Large-scale LLM Training
von: Liang, Mingyu, et al.
Veröffentlicht: (2025)
von: Liang, Mingyu, et al.
Veröffentlicht: (2025)
DIP: Efficient Large Multimodal Model Training with Dynamic Interleaved Pipeline
von: Xue, Zhenliang, et al.
Veröffentlicht: (2025)
von: Xue, Zhenliang, et al.
Veröffentlicht: (2025)
Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices
von: Li, Xiangyu, et al.
Veröffentlicht: (2025)
von: Li, Xiangyu, et al.
Veröffentlicht: (2025)
HETHUB: A Distributed Training System with Heterogeneous Cluster for Large-Scale Models
von: Xu, Si, et al.
Veröffentlicht: (2024)
von: Xu, Si, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Collaborative Split Federated Learning with Parallel Training and Aggregation
von: Papageorgiou, Yiannis, et al.
Veröffentlicht: (2025) -
Accelerating Long-Tail Generation in Synchronous RLHF Training via Adaptive Tensor Parallelism
von: Zhao, Long, et al.
Veröffentlicht: (2026) -
TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training
von: Liu, Man, et al.
Veröffentlicht: (2026) -
FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs
von: Zhang, Haijun, et al.
Veröffentlicht: (2025) -
InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training
von: Wang, Shiju, et al.
Veröffentlicht: (2025)