Don't Let a Few Network Failures Slow the Entire AllReduce
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Peiqing, Jiang, Jiedong, Yu, Nengneng, Wang, Yuefeng, Xiong, Sixian, Wang, Wei, Liu, Zaoxing |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reliable and Resilient Collective Communication Library for LLM Training and Serving
by: Wang, Wei, et al.
Published: (2025)
by: Wang, Wei, et al.
Published: (2025)
Trivance: Latency-Optimal AllReduce by Shortcutting Multiport Networks
by: Juerss, Anton, et al.
Published: (2026)
by: Juerss, Anton, et al.
Published: (2026)
AllReduce Scheduling with Hierarchical Deep Reinforcement Learning
by: Wei, Yufan, et al.
Published: (2025)
by: Wei, Yufan, et al.
Published: (2025)
Short-circuiting Rings for Low-Latency AllReduce
by: Hammer, Sarah-Michelle, et al.
Published: (2025)
by: Hammer, Sarah-Michelle, et al.
Published: (2025)
OptiReduce: Resilient and Tail-Optimal AllReduce for Distributed Deep Learning in the Cloud
by: Warraich, Ertza, et al.
Published: (2023)
by: Warraich, Ertza, et al.
Published: (2023)
Rina: Enhancing Ring-AllReduce with In-network Aggregation in Distributed Model Training
by: Chen, Zixuan, et al.
Published: (2024)
by: Chen, Zixuan, et al.
Published: (2024)
Meili: Enabling SmartNIC as a Service in the Cloud
by: Su, Qiang, et al.
Published: (2023)
by: Su, Qiang, et al.
Published: (2023)
Revisiting the Time Cost Model of AllReduce
by: Xiong, Dian, et al.
Published: (2024)
by: Xiong, Dian, et al.
Published: (2024)
Revisiting Bruck: Phase-Efficient All-to-All Communication in Reconfigurable Networks
by: Juerss, Anton, et al.
Published: (2026)
by: Juerss, Anton, et al.
Published: (2026)
D-LoRa: a Distributed Parameter Adaptation Scheme for LoRa Network
by: Wang, Ruiqi, et al.
Published: (2025)
by: Wang, Ruiqi, et al.
Published: (2025)
Varuna: Enabling Failure-Type Aware RDMA Failover
by: Wang, Xiaoyang, et al.
Published: (2026)
by: Wang, Xiaoyang, et al.
Published: (2026)
RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training
by: Xu, Heng, et al.
Published: (2025)
by: Xu, Heng, et al.
Published: (2025)
FAST: An Efficient Scheduler for All-to-All GPU Communication
by: Lei, Yiran, et al.
Published: (2025)
by: Lei, Yiran, et al.
Published: (2025)
Efficient All-to-All Collective Communication Schedules for Direct-Connect Topologies
by: Basu, Prithwish, et al.
Published: (2023)
by: Basu, Prithwish, et al.
Published: (2023)
Dynamic Hierarchical Birkhoff-von Neumann Decomposition for All-to-All GPU Communication
by: Wu, Yen-Chieh, et al.
Published: (2026)
by: Wu, Yen-Chieh, et al.
Published: (2026)
FODT: Fast, Online, Distributed and Temporary Failure Recovery Approach for MEC
by: Yuan, Xin, et al.
Published: (2023)
by: Yuan, Xin, et al.
Published: (2023)
Fast Multichannel Topology Discovery in Cognitive Radio Networks
by: Wang, Yung-Li, et al.
Published: (2025)
by: Wang, Yung-Li, et al.
Published: (2025)
Q-adaptive: A Multi-Agent Reinforcement Learning Based Routing on Dragonfly Network
by: Kang, Yao, et al.
Published: (2024)
by: Kang, Yao, et al.
Published: (2024)
Multi-Source Coflow Scheduling in Collaborative Edge Computing with Multihop Network
by: Sahni, Yuvraj, et al.
Published: (2024)
by: Sahni, Yuvraj, et al.
Published: (2024)
Towards Timely Video Analytics Services at the Network Edge
by: Li, Xishuo, et al.
Published: (2024)
by: Li, Xishuo, et al.
Published: (2024)
OrchestrRL: Dynamic Compute and Network Orchestration for Disaggregated RL
by: Tan, Xin, et al.
Published: (2026)
by: Tan, Xin, et al.
Published: (2026)
Toward Edge General Intelligence with Multiple-Large Language Model (Multi-LLM): Architecture, Trust, and Orchestration
by: Luo, Haoxiang, et al.
Published: (2025)
by: Luo, Haoxiang, et al.
Published: (2025)
Recursive Offloading for LLM Serving in Multi-tier Networks
by: Wu, Zhiyuan, et al.
Published: (2025)
by: Wu, Zhiyuan, et al.
Published: (2025)
Towards Practical Overlay Networks for Decentralized Federated Learning
by: Hua, Yifan, et al.
Published: (2024)
by: Hua, Yifan, et al.
Published: (2024)
BlockSDN: Towards a High-Performance Blockchain via Software-Defined Cross Networking optimization
by: Jia, Wenyang, et al.
Published: (2025)
by: Jia, Wenyang, et al.
Published: (2025)
PeerSync: Accelerating Containerized Service Delivery at the Network Edge
by: Deng, Yinuo, et al.
Published: (2025)
by: Deng, Yinuo, et al.
Published: (2025)
Arcturus: A Cloud Overlay Network for Global Accelerator with Enhanced Performance and Stability
by: Liu, Matthew Yang, et al.
Published: (2025)
by: Liu, Matthew Yang, et al.
Published: (2025)
On Effectiveness of Graph Neural Network Architectures for Network Digital Twins (NDTs)
by: Zacarias, Iulisloi, et al.
Published: (2025)
by: Zacarias, Iulisloi, et al.
Published: (2025)
Extreme-Scale Interconnection Networks
by: Cano, Alejandro, et al.
Published: (2026)
by: Cano, Alejandro, et al.
Published: (2026)
Throughput-Optimized Networks at Scale
by: Green, Conor James, et al.
Published: (2026)
by: Green, Conor James, et al.
Published: (2026)
Swarm Network-as-a-Service (SNaaS)
by: Alkouz, Balsam, et al.
Published: (2026)
by: Alkouz, Balsam, et al.
Published: (2026)
SPARC-LoRa: A Scalable, Power-efficient, Affordable, Reliable, and Cloud Service-enabled LoRa Networking System for Agriculture Applications
by: Wang, Xi, et al.
Published: (2024)
by: Wang, Xi, et al.
Published: (2024)
Runtime Verification Containers for Publish/Subscribe Networks
by: Mehran, Ali, et al.
Published: (2024)
by: Mehran, Ali, et al.
Published: (2024)
Placing Timely Refreshing Services at the Network Edge
by: Li, Xishuo, et al.
Published: (2024)
by: Li, Xishuo, et al.
Published: (2024)
Revolutionizing Datacenter Networks via Reconfigurable Topologies
by: Avin, Chen, et al.
Published: (2025)
by: Avin, Chen, et al.
Published: (2025)
Don't Wait to be Breached! Creating Asymmetric Uncertainty of Cloud Applications via Moving Target Defenses
by: Torkura, Kennedy A., et al.
Published: (2019)
by: Torkura, Kennedy A., et al.
Published: (2019)
PSMOA: Policy Support Multi-Objective Optimization Algorithm for Decentralized Data Replication
by: Wang, Xi, et al.
Published: (2025)
by: Wang, Xi, et al.
Published: (2025)
RailX: A Flexible, Scalable, and Low-Cost Network Architecture for Hyper-Scale LLM Training Systems
by: Feng, Yinxiao, et al.
Published: (2025)
by: Feng, Yinxiao, et al.
Published: (2025)
SCRamble: Adaptive Decentralized Overlay Construction for Blockchain Networks
by: Kolyvas, Evangelos, et al.
Published: (2026)
by: Kolyvas, Evangelos, et al.
Published: (2026)
Network Anomaly Detection in Distributed Edge Computing Infrastructure
by: Marfo, William, et al.
Published: (2025)
by: Marfo, William, et al.
Published: (2025)
Similar Items
-
Reliable and Resilient Collective Communication Library for LLM Training and Serving
by: Wang, Wei, et al.
Published: (2025) -
Trivance: Latency-Optimal AllReduce by Shortcutting Multiport Networks
by: Juerss, Anton, et al.
Published: (2026) -
AllReduce Scheduling with Hierarchical Deep Reinforcement Learning
by: Wei, Yufan, et al.
Published: (2025) -
Short-circuiting Rings for Low-Latency AllReduce
by: Hammer, Sarah-Michelle, et al.
Published: (2025) -
OptiReduce: Resilient and Tail-Optimal AllReduce for Distributed Deep Learning in the Cloud
by: Warraich, Ertza, et al.
Published: (2023)