Communication-Efficient Large-Scale Distributed Deep Learning: A Comprehensive Survey
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liang, Feng, Zhang, Zhen, Lu, Haifeng, Leung, Victor C. M., Guo, Yanyi, Hu, Xiping |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Resource Allocation and Workload Scheduling for Large-Scale Distributed Deep Learning: A Survey
von: Liang, Feng, et al.
Veröffentlicht: (2024)
von: Liang, Feng, et al.
Veröffentlicht: (2024)
Edge-Cloud Collaborative Computing on Distributed Intelligence and Model Optimization: A Survey
von: Liu, Jing, et al.
Veröffentlicht: (2025)
von: Liu, Jing, et al.
Veröffentlicht: (2025)
PacTrain: Pruning and Adaptive Sparse Gradient Compression for Efficient Collective Communication in Distributed Deep Learning
von: Wang, Yisu, et al.
Veröffentlicht: (2025)
von: Wang, Yisu, et al.
Veröffentlicht: (2025)
FalconFS: Distributed File System for Large-Scale Deep Learning Pipeline
von: Xu, Jingwei, et al.
Veröffentlicht: (2025)
von: Xu, Jingwei, et al.
Veröffentlicht: (2025)
Threats and Defenses in Federated Learning Life Cycle: A Comprehensive Survey and Challenges
von: Li, Yanli, et al.
Veröffentlicht: (2024)
von: Li, Yanli, et al.
Veröffentlicht: (2024)
RapidGNN: Communication Efficient Large-Scale Distributed Training of Graph Neural Networks
von: Niam, Arefin, et al.
Veröffentlicht: (2025)
von: Niam, Arefin, et al.
Veröffentlicht: (2025)
Efficient Training of Large Language Models on Distributed Infrastructures: A Survey
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
von: Hu, Yi-Xiang, et al.
Veröffentlicht: (2026)
von: Hu, Yi-Xiang, et al.
Veröffentlicht: (2026)
Communication-Efficient Distributed Learning via Sparse and Adaptive Stochastic Gradient
von: Deng, Xiaoge, et al.
Veröffentlicht: (2021)
von: Deng, Xiaoge, et al.
Veröffentlicht: (2021)
Entanglement-Efficient Distribution of Quantum Circuits over Large-Scale Quantum Networks
von: Burt, Felix, et al.
Veröffentlicht: (2025)
von: Burt, Felix, et al.
Veröffentlicht: (2025)
Resource-Efficient Compilation of Distributed Quantum Circuits for Solving Large-Scale Wireless Communication Network Problems
von: Chen, Kuan-Cheng, et al.
Veröffentlicht: (2025)
von: Chen, Kuan-Cheng, et al.
Veröffentlicht: (2025)
MPI-Q: A Message Communication Library for Large-Scale Classical-Quantum Heterogeneous Hybrid Distributed Computing
von: Wang, Feng, et al.
Veröffentlicht: (2026)
von: Wang, Feng, et al.
Veröffentlicht: (2026)
MegaScale-Data: Scaling Dataloader for Multisource Large Foundation Model Training
von: Zhao, Juntao, et al.
Veröffentlicht: (2025)
von: Zhao, Juntao, et al.
Veröffentlicht: (2025)
ACE-Sync: An Adaptive Cloud-Edge Synchronization Framework for Communication-Efficient Large-Scale Distributed Model Training
von: Yang, Yi, et al.
Veröffentlicht: (2025)
von: Yang, Yi, et al.
Veröffentlicht: (2025)
SparDL: Distributed Deep Learning Training with Efficient Sparse Communication
von: Zhao, Minjun, et al.
Veröffentlicht: (2023)
von: Zhao, Minjun, et al.
Veröffentlicht: (2023)
Lumos: Efficient Performance Modeling and Estimation for Large-scale LLM Training
von: Liang, Mingyu, et al.
Veröffentlicht: (2025)
von: Liang, Mingyu, et al.
Veröffentlicht: (2025)
Byzantine-Robust and Communication-Efficient Distributed Training: Compressive and Cyclic Gradient Coding
von: Li, Chengxi, et al.
Veröffentlicht: (2026)
von: Li, Chengxi, et al.
Veröffentlicht: (2026)
HETHUB: A Distributed Training System with Heterogeneous Cluster for Large-Scale Models
von: Xu, Si, et al.
Veröffentlicht: (2024)
von: Xu, Si, et al.
Veröffentlicht: (2024)
Mesh-Attention: A New Communication-Efficient Distributed Attention with Improved Data Locality
von: Chen, Sirui, et al.
Veröffentlicht: (2025)
von: Chen, Sirui, et al.
Veröffentlicht: (2025)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
von: Zhang, WenZheng, et al.
Veröffentlicht: (2024)
von: Zhang, WenZheng, et al.
Veröffentlicht: (2024)
COMET: A Comprehensive Cluster Design Methodology for Distributed Deep Learning Training
von: Kadiyala, Divya Kiran, et al.
Veröffentlicht: (2022)
von: Kadiyala, Divya Kiran, et al.
Veröffentlicht: (2022)
Optimizing High-Throughput Distributed Data Pipelines for Reproducible Deep Learning at Scale
von: Mittal, Kashish, et al.
Veröffentlicht: (2026)
von: Mittal, Kashish, et al.
Veröffentlicht: (2026)
DeepServe: Serverless Large Language Model Serving at Scale
von: Hu, Junhao, et al.
Veröffentlicht: (2025)
von: Hu, Junhao, et al.
Veröffentlicht: (2025)
Verify Distributed Deep Learning Model Implementation Refinement with Iterative Relation Inference
von: Wang, Zhanghan, et al.
Veröffentlicht: (2025)
von: Wang, Zhanghan, et al.
Veröffentlicht: (2025)
High-Dimensional Data Processing: Benchmarking Machine Learning and Deep Learning Architectures in Local and Distributed Environments
von: Rodriguez, Julian, et al.
Veröffentlicht: (2025)
von: Rodriguez, Julian, et al.
Veröffentlicht: (2025)
Acceleration for Deep Reinforcement Learning using Parallel and Distributed Computing: A Survey
von: Liu, Zhihong, et al.
Veröffentlicht: (2024)
von: Liu, Zhihong, et al.
Veröffentlicht: (2024)
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models
von: Cheng, Jialiang, et al.
Veröffentlicht: (2024)
von: Cheng, Jialiang, et al.
Veröffentlicht: (2024)
Mist: Efficient Distributed Training of Large Language Models via Memory-Parallelism Co-Optimization
von: Zhu, Zhanda, et al.
Veröffentlicht: (2025)
von: Zhu, Zhanda, et al.
Veröffentlicht: (2025)
MegatronApp: Efficient and Comprehensive Management on Distributed LLM Training
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
ScaleSim: Serving Large-Scale Multi-Agent Simulation with Invocation Distance-Based Memory Management
von: Pan, Zaifeng, et al.
Veröffentlicht: (2026)
von: Pan, Zaifeng, et al.
Veröffentlicht: (2026)
Placement Semantics for Distributed Deep Learning: A Systematic Framework for Analyzing Parallelism Strategies
von: Mehta, Deep Pankajbhai
Veröffentlicht: (2026)
von: Mehta, Deep Pankajbhai
Veröffentlicht: (2026)
Lagom: Unleashing the Power of Communication and Computation Overlapping for Distributed LLM Training
von: Xu, Guanbin, et al.
Veröffentlicht: (2026)
von: Xu, Guanbin, et al.
Veröffentlicht: (2026)
Learning Provably Correct Distributed Protocols Without Human Knowledge
von: Hui, Yujie, et al.
Veröffentlicht: (2026)
von: Hui, Yujie, et al.
Veröffentlicht: (2026)
A Survey on Large Language Model Acceleration based on KV Cache Management
von: Li, Haoyang, et al.
Veröffentlicht: (2024)
von: Li, Haoyang, et al.
Veröffentlicht: (2024)
Communication-Efficient Distributed Deep Learning via Federated Dynamic Averaging
von: Theologitis, Michail, et al.
Veröffentlicht: (2024)
von: Theologitis, Michail, et al.
Veröffentlicht: (2024)
Towards Carbon-Aware Container Orchestration: Predicting Workload Energy Consumption with Federated Learning
von: Saad, Zainab, et al.
Veröffentlicht: (2025)
von: Saad, Zainab, et al.
Veröffentlicht: (2025)
Demystifying the Communication Characteristics for Distributed Transformer Models
von: Anthony, Quentin, et al.
Veröffentlicht: (2024)
von: Anthony, Quentin, et al.
Veröffentlicht: (2024)
Characterizing Compute-Communication Overlap in GPU-Accelerated Distributed Deep Learning: Performance and Power Implications
von: Lee, Seonho, et al.
Veröffentlicht: (2025)
von: Lee, Seonho, et al.
Veröffentlicht: (2025)
A Survey of Synchronization Technologies for Low-power Backscatter Communication
von: Jiang, Wenyuan, et al.
Veröffentlicht: (2025)
von: Jiang, Wenyuan, et al.
Veröffentlicht: (2025)
MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services
von: Yu, Dianhai, et al.
Veröffentlicht: (2022)
von: Yu, Dianhai, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
Resource Allocation and Workload Scheduling for Large-Scale Distributed Deep Learning: A Survey
von: Liang, Feng, et al.
Veröffentlicht: (2024) -
Edge-Cloud Collaborative Computing on Distributed Intelligence and Model Optimization: A Survey
von: Liu, Jing, et al.
Veröffentlicht: (2025) -
PacTrain: Pruning and Adaptive Sparse Gradient Compression for Efficient Collective Communication in Distributed Deep Learning
von: Wang, Yisu, et al.
Veröffentlicht: (2025) -
FalconFS: Distributed File System for Large-Scale Deep Learning Pipeline
von: Xu, Jingwei, et al.
Veröffentlicht: (2025) -
Threats and Defenses in Federated Learning Life Cycle: A Comprehensive Survey and Challenges
von: Li, Yanli, et al.
Veröffentlicht: (2024)