RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Heng, Yu, Zhiwei, Du, Chengze, Zhou, Ying, Li, Letian, Wang, Haojie, Cheng, Weiqiang, Li, Jialong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Temporal-Aware GPU Resource Allocation for Distributed LLM Inference via Reinforcement Learning
by: Du, Chengze, et al.
Published: (2025)
by: Du, Chengze, et al.
Published: (2025)
FAST: An Efficient Scheduler for All-to-All GPU Communication
by: Lei, Yiran, et al.
Published: (2025)
by: Lei, Yiran, et al.
Published: (2025)
Ethereal: Divide and Conquer Network Load Balancing in Large-Scale Distributed Training
by: Addanki, Vamsi, et al.
Published: (2024)
by: Addanki, Vamsi, et al.
Published: (2024)
Efficient All-to-All Collective Communication Schedules for Direct-Connect Topologies
by: Basu, Prithwish, et al.
Published: (2023)
by: Basu, Prithwish, et al.
Published: (2023)
Revisiting Bruck: Phase-Efficient All-to-All Communication in Reconfigurable Networks
by: Juerss, Anton, et al.
Published: (2026)
by: Juerss, Anton, et al.
Published: (2026)
Dynamic Hierarchical Birkhoff-von Neumann Decomposition for All-to-All GPU Communication
by: Wu, Yen-Chieh, et al.
Published: (2026)
by: Wu, Yen-Chieh, et al.
Published: (2026)
RailX: A Flexible, Scalable, and Low-Cost Network Architecture for Hyper-Scale LLM Training Systems
by: Feng, Yinxiao, et al.
Published: (2025)
by: Feng, Yinxiao, et al.
Published: (2025)
Rina: Enhancing Ring-AllReduce with In-network Aggregation in Distributed Model Training
by: Chen, Zixuan, et al.
Published: (2024)
by: Chen, Zixuan, et al.
Published: (2024)
OptiReduce: Resilient and Tail-Optimal AllReduce for Distributed Deep Learning in the Cloud
by: Warraich, Ertza, et al.
Published: (2023)
by: Warraich, Ertza, et al.
Published: (2023)
Deduplicator: When Computation Reuse Meets Load Balancing at the Network Edge
by: Azad, Md Washik Al, et al.
Published: (2024)
by: Azad, Md Washik Al, et al.
Published: (2024)
AllReduce Scheduling with Hierarchical Deep Reinforcement Learning
by: Wei, Yufan, et al.
Published: (2025)
by: Wei, Yufan, et al.
Published: (2025)
Short-circuiting Rings for Low-Latency AllReduce
by: Hammer, Sarah-Michelle, et al.
Published: (2025)
by: Hammer, Sarah-Michelle, et al.
Published: (2025)
Optimal Oblivious Load-Balancing for Sparse Traffic in Large-Scale Satellite Networks
by: Ramakanth, Rudrapatna Vallabh, et al.
Published: (2026)
by: Ramakanth, Rudrapatna Vallabh, et al.
Published: (2026)
QoS-Aware Load Balancing in the Computing Continuum via Multi-Player Bandits
by: Čilić, Ivan, et al.
Published: (2025)
by: Čilić, Ivan, et al.
Published: (2025)
Trivance: Latency-Optimal AllReduce by Shortcutting Multiport Networks
by: Juerss, Anton, et al.
Published: (2026)
by: Juerss, Anton, et al.
Published: (2026)
SlimCaching: Edge Caching of Mixture-of-Experts for Distributed Inference
by: Chen, Qian, et al.
Published: (2025)
by: Chen, Qian, et al.
Published: (2025)
GORGO: Maximizing KV-Cache Reuse While Minimizing Network Latency in Cross-Region LLM Load Balancing
by: Toniolo, Alessio Ricci, et al.
Published: (2026)
by: Toniolo, Alessio Ricci, et al.
Published: (2026)
SpaceMoE: Realizing Distributed Mixture-of-Experts Inference over Space Networks
by: Wang, Zhanwei, et al.
Published: (2026)
by: Wang, Zhanwei, et al.
Published: (2026)
Heterogeneity-aware P2P Wireless Energy Transfer for Balanced Energy Distribution
by: Ojha, Tamoghna, et al.
Published: (2022)
by: Ojha, Tamoghna, et al.
Published: (2022)
FODT: Fast, Online, Distributed and Temporary Failure Recovery Approach for MEC
by: Yuan, Xin, et al.
Published: (2023)
by: Yuan, Xin, et al.
Published: (2023)
SANSee: A Physical-layer Semantic-aware Networking Framework for Distributed Wireless Sensing
by: Zhu, Huixiang, et al.
Published: (2024)
by: Zhu, Huixiang, et al.
Published: (2024)
A Survey on Resource Management in Joint Communication and Computing-Embedded SAGIN
by: Chen, Qian, et al.
Published: (2024)
by: Chen, Qian, et al.
Published: (2024)
LIMO: Load-balanced Offloading with MAPE and Particle Swarm Optimization in Mobile Fog Networks
by: Seraj, Yasaman, et al.
Published: (2024)
by: Seraj, Yasaman, et al.
Published: (2024)
From Skew to Symmetry: Node-Interconnect Multi-Path Balancing with Execution-time Planning for Modern GPU Clusters
by: Yao, Jinghan, et al.
Published: (2026)
by: Yao, Jinghan, et al.
Published: (2026)
Don't Let a Few Network Failures Slow the Entire AllReduce
by: Chen, Peiqing, et al.
Published: (2026)
by: Chen, Peiqing, et al.
Published: (2026)
cMPI: Using CXL Memory Sharing for MPI One-Sided and Two-Sided Inter-Node Communications
by: Wang, Xi, et al.
Published: (2025)
by: Wang, Xi, et al.
Published: (2025)
Enabling Scalability in Asynchronous and Bidirectional Communication in LPWAN
by: Rahman, Mahbubur
Published: (2025)
by: Rahman, Mahbubur
Published: (2025)
Network Anomaly Detection in Distributed Edge Computing Infrastructure
by: Marfo, William, et al.
Published: (2025)
by: Marfo, William, et al.
Published: (2025)
A Uniqueness Theorem for Distributed Computation under Physical Constraint
by: Ren, Zhiyuan, et al.
Published: (2025)
by: Ren, Zhiyuan, et al.
Published: (2025)
Policy Design in Zero-Trust Distributed Networks: Challenges and Solutions
by: Sandjaja, Fannya R., et al.
Published: (2025)
by: Sandjaja, Fannya R., et al.
Published: (2025)
DistriFS: A Platform and User Agnostic Approach to File Distribution
by: Boesch, Julian
Published: (2024)
by: Boesch, Julian
Published: (2024)
Fast Prototyping of Distributed Stream Processing Applications with stream2gym
by: Ifath, Md. Monzurul Amin, et al.
Published: (2024)
by: Ifath, Md. Monzurul Amin, et al.
Published: (2024)
A Multi-Layered Distributed Computing Framework for Enhanced Edge Computing
by: Ma, Ke, et al.
Published: (2024)
by: Ma, Ke, et al.
Published: (2024)
FlowTracer: A Tool for Uncovering Network Path Usage Imbalance in AI Training Clusters
by: Jamil, Hasibul, et al.
Published: (2024)
by: Jamil, Hasibul, et al.
Published: (2024)
ScaleAcross Explorer: Exploring Communication Optimization for Scale-Across AI Model Training
by: Li, Minghao, et al.
Published: (2026)
by: Li, Minghao, et al.
Published: (2026)
Harvest: Adaptive Photonic Switching Schedules for Collective Communication in Scale-up Domains
by: Rahman, Mahir, et al.
Published: (2026)
by: Rahman, Mahir, et al.
Published: (2026)
OptiNIC: A Resilient and Tail-Optimal RDMA NIC for Distributed ML Workloads
by: Warraich, Ertza, et al.
Published: (2025)
by: Warraich, Ertza, et al.
Published: (2025)
D-LoRa: a Distributed Parameter Adaptation Scheme for LoRa Network
by: Wang, Ruiqi, et al.
Published: (2025)
by: Wang, Ruiqi, et al.
Published: (2025)
LOAM: Low-latency Communication, Caching, and Computation Placement in Data-Intensive Computing Networks
by: Zhang, Jinkun, et al.
Published: (2024)
by: Zhang, Jinkun, et al.
Published: (2024)
A Density-Delay Law for Stable Event-Driven State Progression in Open Distributed Systems
by: Chen, Bin, et al.
Published: (2026)
by: Chen, Bin, et al.
Published: (2026)
Similar Items
-
Temporal-Aware GPU Resource Allocation for Distributed LLM Inference via Reinforcement Learning
by: Du, Chengze, et al.
Published: (2025) -
FAST: An Efficient Scheduler for All-to-All GPU Communication
by: Lei, Yiran, et al.
Published: (2025) -
Ethereal: Divide and Conquer Network Load Balancing in Large-Scale Distributed Training
by: Addanki, Vamsi, et al.
Published: (2024) -
Efficient All-to-All Collective Communication Schedules for Direct-Connect Topologies
by: Basu, Prithwish, et al.
Published: (2023) -
Revisiting Bruck: Phase-Efficient All-to-All Communication in Reconfigurable Networks
by: Juerss, Anton, et al.
Published: (2026)