Communication Optimization for Distributed Training: Architecture, Advances, and Opportunities
Fuente:
arXiv
Saved in:
| Main Authors: | Wei, Yunze, Hu, Tianshuo, Liang, Cong, Cui, Yong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Robust Fully-Asynchronous Methods for Distributed Training over General Architecture
by: Zhu, Zehan, et al.
Published: (2023)
by: Zhu, Zehan, et al.
Published: (2023)
Hybrid Dual-Batch and Cyclic Progressive Learning for Efficient Distributed Training
by: Lu, Kuan-Wei, et al.
Published: (2025)
by: Lu, Kuan-Wei, et al.
Published: (2025)
Low-Communication Resilient Distributed Estimation Algorithm Based on Memory Mechanism
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
Spindle: Efficient Distributed Training of Multi-Task Large Models via Wavefront Scheduling
by: Wang, Yujie, et al.
Published: (2024)
by: Wang, Yujie, et al.
Published: (2024)
Communication-Efficient Federated Group Distributionally Robust Optimization
by: Guo, Zhishuai, et al.
Published: (2024)
by: Guo, Zhishuai, et al.
Published: (2024)
SignMuon: Communication-Efficient Distributed Muon Optimization
by: Mishra, Neel, et al.
Published: (2026)
by: Mishra, Neel, et al.
Published: (2026)
SparDL: Distributed Deep Learning Training with Efficient Sparse Communication
by: Zhao, Minjun, et al.
Published: (2023)
by: Zhao, Minjun, et al.
Published: (2023)
Communication-Efficient Distributed Training for Collaborative Flat Optima Recovery in Deep Learning
by: Dimlioglu, Tolga, et al.
Published: (2025)
by: Dimlioglu, Tolga, et al.
Published: (2025)
Distributed Low-Communication Training with Decoupled Momentum Optimization
by: Nedelkoski, Sasho, et al.
Published: (2025)
by: Nedelkoski, Sasho, et al.
Published: (2025)
Echo: Simulating Distributed Training At Scale
by: Feng, Yicheng, et al.
Published: (2024)
by: Feng, Yicheng, et al.
Published: (2024)
OpenDiLoCo: An Open-Source Framework for Globally Distributed Low-Communication Training
by: Jaghouar, Sami, et al.
Published: (2024)
by: Jaghouar, Sami, et al.
Published: (2024)
Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training and Inference
by: Luo, Shuqing, et al.
Published: (2025)
by: Luo, Shuqing, et al.
Published: (2025)
SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM Training
by: Jia, Jinda, et al.
Published: (2024)
by: Jia, Jinda, et al.
Published: (2024)
Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing
by: Liu, Mengfan, et al.
Published: (2025)
by: Liu, Mengfan, et al.
Published: (2025)
CDFGNN: a Systematic Design of Cache-based Distributed Full-Batch Graph Neural Network Training with Communication Reduction
by: Zhang, Shuai, et al.
Published: (2024)
by: Zhang, Shuai, et al.
Published: (2024)
Distributed Training under Packet Loss
by: Weintraub, Erez, et al.
Published: (2025)
by: Weintraub, Erez, et al.
Published: (2025)
LIBRA: Enabling Workload-aware Multi-dimensional Network Topology Optimization for Distributed Training of Large AI Models
by: Won, William, et al.
Published: (2021)
by: Won, William, et al.
Published: (2021)
Incentivised Orchestrated Training Architecture (IOTA): A Technical Primer for Release
by: Quinque, Felix, et al.
Published: (2025)
by: Quinque, Felix, et al.
Published: (2025)
AntDT: A Self-Adaptive Distributed Training Framework for Leader and Straggler Nodes
by: Xiao, Youshao, et al.
Published: (2024)
by: Xiao, Youshao, et al.
Published: (2024)
On Optimizing the Communication of Model Parallelism
by: Zhuang, Yonghao, et al.
Published: (2022)
by: Zhuang, Yonghao, et al.
Published: (2022)
MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production
by: Jin, Chao, et al.
Published: (2025)
by: Jin, Chao, et al.
Published: (2025)
Empowering Distributed Training with Sparsity-driven Data Synchronization
by: Wang, Zhuang, et al.
Published: (2023)
by: Wang, Zhuang, et al.
Published: (2023)
Accelerating Wireless Distributed Learning via Hybrid Split and Federated Learning Optimization
by: Guo, Kun, et al.
Published: (2025)
by: Guo, Kun, et al.
Published: (2025)
Distributed Convolutional Neural Network Training on Mobile and Edge Clusters
by: Rama, Pranav, et al.
Published: (2024)
by: Rama, Pranav, et al.
Published: (2024)
Efficient Parallelization Layouts for Large-Scale Distributed Model Training
by: Hagemann, Johannes, et al.
Published: (2023)
by: Hagemann, Johannes, et al.
Published: (2023)
Lion Cub: Minimizing Communication Overhead in Distributed Lion
by: Ishikawa, Satoki, et al.
Published: (2024)
by: Ishikawa, Satoki, et al.
Published: (2024)
AI-in-the-Loop Sensing and Communication Joint Design for Edge Intelligence
by: Cai, Zhijie, et al.
Published: (2025)
by: Cai, Zhijie, et al.
Published: (2025)
Towards Communication-efficient Federated Learning via Sparse and Aligned Adaptive Optimization
by: Deng, Xiumei, et al.
Published: (2024)
by: Deng, Xiumei, et al.
Published: (2024)
AI-Driven Health Monitoring of Distributed Computing Architecture: Insights from XGBoost and SHAP
by: Sun, Xiaoxuan, et al.
Published: (2024)
by: Sun, Xiaoxuan, et al.
Published: (2024)
Decentralized Orchestration Architecture for Fluid Computing: A Secure Distributed AI Use Case
by: Cajaraville-Aboy, Diego, et al.
Published: (2026)
by: Cajaraville-Aboy, Diego, et al.
Published: (2026)
Federated Communication-Efficient Multi-Objective Optimization
by: Askin, Baris, et al.
Published: (2024)
by: Askin, Baris, et al.
Published: (2024)
Federated Learning Optimization: A Comparative Study of Data and Model Exchange Strategies in Dynamic Networks
by: Luqman, Alka, et al.
Published: (2024)
by: Luqman, Alka, et al.
Published: (2024)
Minder: Faulty Machine Detection for Large-scale Distributed Model Training
by: Deng, Yangtao, et al.
Published: (2024)
by: Deng, Yangtao, et al.
Published: (2024)
QSync: Quantization-Minimized Synchronous Distributed Training Across Hybrid Devices
by: Zhao, Juntao, et al.
Published: (2024)
by: Zhao, Juntao, et al.
Published: (2024)
Fully Distributed Online Training of Graph Neural Networks in Networked Systems
by: Olshevskyi, Rostyslav, et al.
Published: (2024)
by: Olshevskyi, Rostyslav, et al.
Published: (2024)
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training
by: Fernandez, Jared, et al.
Published: (2024)
by: Fernandez, Jared, et al.
Published: (2024)
Leiden-Fusion Partitioning Method for Effective Distributed Training of Graph Embeddings
by: Bai, Yuhe, et al.
Published: (2024)
by: Bai, Yuhe, et al.
Published: (2024)
An Experimental Comparison of Partitioning Strategies for Distributed Graph Neural Network Training
by: Merkel, Nikolai, et al.
Published: (2023)
by: Merkel, Nikolai, et al.
Published: (2023)
Characterizing the Efficiency of Distributed Training: A Power, Performance, and Thermal Perspective
by: Go, Seokjin, et al.
Published: (2025)
by: Go, Seokjin, et al.
Published: (2025)
Communication-Efficient Distributed Learning with Local Immediate Error Compensation
by: Cheng, Yifei, et al.
Published: (2024)
by: Cheng, Yifei, et al.
Published: (2024)
Similar Items
-
Robust Fully-Asynchronous Methods for Distributed Training over General Architecture
by: Zhu, Zehan, et al.
Published: (2023) -
Hybrid Dual-Batch and Cyclic Progressive Learning for Efficient Distributed Training
by: Lu, Kuan-Wei, et al.
Published: (2025) -
Low-Communication Resilient Distributed Estimation Algorithm Based on Memory Mechanism
by: Li, Wei, et al.
Published: (2025) -
Spindle: Efficient Distributed Training of Multi-Task Large Models via Wavefront Scheduling
by: Wang, Yujie, et al.
Published: (2024) -
Communication-Efficient Federated Group Distributionally Robust Optimization
by: Guo, Zhishuai, et al.
Published: (2024)