Efficient AllReduce with Stragglers
Fuente:
arXiv
Saved in:
| Main Authors: | Devraj, Arjun, Ding, Eric, Kumar, Abhishek Vijaya, Kleinberg, Robert, Singh, Rachee |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PCCL: Photonic circuit-switched collective communication for distributed ML
by: Kumar, Abhishek Vijaya, et al.
Published: (2025)
by: Kumar, Abhishek Vijaya, et al.
Published: (2025)
Short-circuiting Rings for Low-Latency AllReduce
by: Hammer, Sarah-Michelle, et al.
Published: (2025)
by: Hammer, Sarah-Michelle, et al.
Published: (2025)
AQUA: Network-Accelerated Memory Offloading for LLMs in Scale-Up GPU Domains
by: Kumar, Abhishek Vijaya, et al.
Published: (2024)
by: Kumar, Abhishek Vijaya, et al.
Published: (2024)
Revisiting the Time Cost Model of AllReduce
by: Xiong, Dian, et al.
Published: (2024)
by: Xiong, Dian, et al.
Published: (2024)
Don't Let a Few Network Failures Slow the Entire AllReduce
by: Chen, Peiqing, et al.
Published: (2026)
by: Chen, Peiqing, et al.
Published: (2026)
CCL-Bench 1.0: A Trace-Based Benchmark for LLM Infrastructure
by: Ding, Eric, et al.
Published: (2026)
by: Ding, Eric, et al.
Published: (2026)
General Coded Computing in a Probabilistic Straggler Regime
by: Moradi, Parsa, et al.
Published: (2025)
by: Moradi, Parsa, et al.
Published: (2025)
Understanding Stragglers in Large Model Training Using What-if Analysis
by: Lin, Jinkun, et al.
Published: (2025)
by: Lin, Jinkun, et al.
Published: (2025)
Straggler-Resilient Decentralized Learning via Adaptive Asynchronous Updates
by: Xiong, Guojun, et al.
Published: (2023)
by: Xiong, Guojun, et al.
Published: (2023)
AllReduce Scheduling with Hierarchical Deep Reinforcement Learning
by: Wei, Yufan, et al.
Published: (2025)
by: Wei, Yufan, et al.
Published: (2025)
Distributed Learning based on 1-Bit Gradient Coding in the Presence of Stragglers
by: Li, Chengxi, et al.
Published: (2024)
by: Li, Chengxi, et al.
Published: (2024)
FlashMoE: Fast Distributed MoE in a Single Kernel
by: Aimuyo, Osayamen Jonathan, et al.
Published: (2025)
by: Aimuyo, Osayamen Jonathan, et al.
Published: (2025)
Stragglers Can Contribute More: Uncertainty-Aware Distillation for Asynchronous Federated Learning
by: Wang, Yujia, et al.
Published: (2025)
by: Wang, Yujia, et al.
Published: (2025)
Eliminating Hidden Serialization in Multi-Node Megakernel Communication
by: Oh, Byungsoo, et al.
Published: (2026)
by: Oh, Byungsoo, et al.
Published: (2026)
An All-Reduce Compatible Top-K Compressor for Communication-Efficient Distributed Learning
by: Chen, Chuyan, et al.
Published: (2025)
by: Chen, Chuyan, et al.
Published: (2025)
Trivance: Latency-Optimal AllReduce by Shortcutting Multiport Networks
by: Juerss, Anton, et al.
Published: (2026)
by: Juerss, Anton, et al.
Published: (2026)
AntDT: A Self-Adaptive Distributed Training Framework for Leader and Straggler Nodes
by: Xiao, Youshao, et al.
Published: (2024)
by: Xiao, Youshao, et al.
Published: (2024)
OptiReduce: Resilient and Tail-Optimal AllReduce for Distributed Deep Learning in the Cloud
by: Warraich, Ertza, et al.
Published: (2023)
by: Warraich, Ertza, et al.
Published: (2023)
FedCore: Straggler-Free Federated Learning with Distributed Coresets
by: Guo, Hongpeng, et al.
Published: (2024)
by: Guo, Hongpeng, et al.
Published: (2024)
Towards Straggler-Resilient Split Federated Learning: An Unbalanced Update Approach
by: Liang, Dandan, et al.
Published: (2025)
by: Liang, Dandan, et al.
Published: (2025)
CLIP: Client-Side Invariant Pruning for Mitigating Stragglers in Secure Federated Learning
by: DiMaggio, Anthony, et al.
Published: (2025)
by: DiMaggio, Anthony, et al.
Published: (2025)
Guard: Scalable Straggler Detection and Node Health Management for Large-Scale Training
by: Liu, Guanliang, et al.
Published: (2026)
by: Liu, Guanliang, et al.
Published: (2026)
Rina: Enhancing Ring-AllReduce with In-network Aggregation in Distributed Model Training
by: Chen, Zixuan, et al.
Published: (2024)
by: Chen, Zixuan, et al.
Published: (2024)
SignMuon: Communication-Efficient Distributed Muon Optimization
by: Mishra, Neel, et al.
Published: (2026)
by: Mishra, Neel, et al.
Published: (2026)
Hybrid Dual-Batch and Cyclic Progressive Learning for Efficient Distributed Training
by: Lu, Kuan-Wei, et al.
Published: (2025)
by: Lu, Kuan-Wei, et al.
Published: (2025)
Reducing Energy Bloat in Large Model Training
by: Chung, Jae-Won, et al.
Published: (2023)
by: Chung, Jae-Won, et al.
Published: (2023)
All is Not Lost: LLM Recovery without Checkpoints
by: Blagoev, Nikolay, et al.
Published: (2025)
by: Blagoev, Nikolay, et al.
Published: (2025)
Straggler Tolerant and Resilient DL Training on Homogeneous GPUs
by: Zhang, Zeyu, et al.
Published: (2025)
by: Zhang, Zeyu, et al.
Published: (2025)
DrJAX: Scalable and Differentiable MapReduce Primitives in JAX
by: Rush, Keith, et al.
Published: (2024)
by: Rush, Keith, et al.
Published: (2024)
Reducing Communication for Split Learning by Randomized Top-k Sparsification
by: Zheng, Fei, et al.
Published: (2023)
by: Zheng, Fei, et al.
Published: (2023)
Exploiting Stragglers in Distributed Computing Systems with Task Grouping
by: Adikari, Tharindu, et al.
Published: (2024)
by: Adikari, Tharindu, et al.
Published: (2024)
Deal: Distributed End-to-End GNN Inference for All Nodes
by: Chen, Shiyang, et al.
Published: (2025)
by: Chen, Shiyang, et al.
Published: (2025)
Efficient Long-context Language Model Training by Core Attention Disaggregation
by: Zhuang, Yonghao, et al.
Published: (2025)
by: Zhuang, Yonghao, et al.
Published: (2025)
LLMBridge: Reducing Costs to Access LLMs in a Prompt-Centric Internet
by: Martin, Noah, et al.
Published: (2024)
by: Martin, Noah, et al.
Published: (2024)
Parallel Track Transformers: Enabling Fast GPU Inference with Reduced Synchronization
by: Wang, Chong, et al.
Published: (2026)
by: Wang, Chong, et al.
Published: (2026)
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models
by: Ding, Shiwei, et al.
Published: (2025)
by: Ding, Shiwei, et al.
Published: (2025)
Memory and Bandwidth are All You Need for Fully Sharded Data Parallel
by: Wang, Jiangtao, et al.
Published: (2025)
by: Wang, Jiangtao, et al.
Published: (2025)
Reducing Memory Contention and I/O Congestion for Disk-based GNN Training
by: Jiang, Qisheng, et al.
Published: (2024)
by: Jiang, Qisheng, et al.
Published: (2024)
Not All Federated Learning Algorithms Are Created Equal: A Performance Evaluation Study
by: Baumgart, Gustav A., et al.
Published: (2024)
by: Baumgart, Gustav A., et al.
Published: (2024)
Reducing Communication Overhead in Federated Learning for Network Anomaly Detection with Adaptive Client Selection
by: Marfo, William, et al.
Published: (2025)
by: Marfo, William, et al.
Published: (2025)
Similar Items
-
PCCL: Photonic circuit-switched collective communication for distributed ML
by: Kumar, Abhishek Vijaya, et al.
Published: (2025) -
Short-circuiting Rings for Low-Latency AllReduce
by: Hammer, Sarah-Michelle, et al.
Published: (2025) -
AQUA: Network-Accelerated Memory Offloading for LLMs in Scale-Up GPU Domains
by: Kumar, Abhishek Vijaya, et al.
Published: (2024) -
Revisiting the Time Cost Model of AllReduce
by: Xiong, Dian, et al.
Published: (2024) -
Don't Let a Few Network Failures Slow the Entire AllReduce
by: Chen, Peiqing, et al.
Published: (2026)