Near-Optimal Sparse Allreduce for Distributed Deep Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Shigang, Hoefler, Torsten |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AutoDDL: Automatic Distributed Deep Learning with Near-Optimal Bandwidth Cost
by: Chen, Jinfan, et al.
Published: (2023)
by: Chen, Jinfan, et al.
Published: (2023)
Chimera: Efficiently Training Large-Scale Neural Networks with Bidirectional Pipelines
by: Li, Shigang, et al.
Published: (2021)
by: Li, Shigang, et al.
Published: (2021)
FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor Cores
by: Shi, Jinliang, et al.
Published: (2024)
by: Shi, Jinliang, et al.
Published: (2024)
Libra: Unleashing GPU Heterogeneity for High-Performance Sparse Matrix Multiplication
by: Shi, Jinliang, et al.
Published: (2025)
by: Shi, Jinliang, et al.
Published: (2025)
SparkAttention: High-Performance Multi-Head Attention for Large Models on Volta GPU Architecture
by: Xu, Youxuan, et al.
Published: (2025)
by: Xu, Youxuan, et al.
Published: (2025)
Quantize Once, Train Fast: Allreduce-Compatible Compression with Provable Guarantees
by: Xin, Jihao, et al.
Published: (2023)
by: Xin, Jihao, et al.
Published: (2023)
HFedATM: Hierarchical Federated Domain Generalization via Optimal Transport and Regularized Mean Aggregation
by: Nguyen, Thinh, et al.
Published: (2025)
by: Nguyen, Thinh, et al.
Published: (2025)
Impact of Network Topology on Byzantine Resilience in Decentralized Federated Learning
by: Bhattacharya, Siddhartha, et al.
Published: (2024)
by: Bhattacharya, Siddhartha, et al.
Published: (2024)
How Machine Learning-Data Driven Replication Strategies Enhance Fault Tolerance in Large-Scale Distributed Systems
by: Murimi, Almond Kiruthu
Published: (2025)
by: Murimi, Almond Kiruthu
Published: (2025)
Token Coherence: Adapting MESI Cache Protocols to Minimize Synchronization Overhead in Multi-Agent LLM Systems
by: Parakhin, Vladyslav
Published: (2026)
by: Parakhin, Vladyslav
Published: (2026)
Towards Optimal Heterogeneous Client Sampling in Multi-Model Federated Learning
by: Zhang, Haoran, et al.
Published: (2025)
by: Zhang, Haoran, et al.
Published: (2025)
TAGC: Optimizing Gradient Communication in Distributed Transformer Training
by: Polyakov, Igor, et al.
Published: (2025)
by: Polyakov, Igor, et al.
Published: (2025)
Asynchronous Multi-Server Federated Learning for Geo-Distributed Clients
by: Zuo, Yuncong, et al.
Published: (2024)
by: Zuo, Yuncong, et al.
Published: (2024)
SAGA: Workflow-Atomic Scheduling for AI Agent Inference on GPU Clusters
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
GRAIN: Exact Graph Reconstruction from Gradients
by: Drencheva, Maria, et al.
Published: (2025)
by: Drencheva, Maria, et al.
Published: (2025)
Federated Learning Model Aggregation in Heterogenous Aerial and Space Networks
by: Dong, Fan, et al.
Published: (2023)
by: Dong, Fan, et al.
Published: (2023)
AMP4EC: Adaptive Model Partitioning Framework for Efficient Deep Learning Inference in Edge Computing Environments
by: Zhang, Guilin, et al.
Published: (2025)
by: Zhang, Guilin, et al.
Published: (2025)
LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management
by: Xiong, Yi, et al.
Published: (2024)
by: Xiong, Yi, et al.
Published: (2024)
Asynchronous Byzantine Federated Learning
by: Cox, Bart, et al.
Published: (2024)
by: Cox, Bart, et al.
Published: (2024)
Training Diffusion Models with Federated Learning
by: de Goede, Matthijs, et al.
Published: (2024)
by: de Goede, Matthijs, et al.
Published: (2024)
Aergia: Leveraging Heterogeneity in Federated Learning Systems
by: Cox, Bart, et al.
Published: (2022)
by: Cox, Bart, et al.
Published: (2022)
Parameterizing Federated Continual Learning for Reproducible Research
by: Cox, Bart, et al.
Published: (2024)
by: Cox, Bart, et al.
Published: (2024)
Breaking (Global) Barriers in Parallel Stochastic Optimization with Wait-Avoiding Group Averaging
by: Li, Shigang, et al.
Published: (2020)
by: Li, Shigang, et al.
Published: (2020)
Hyper-parameter Optimization for Federated Learning with Step-wise Adaptive Mechanism
by: Saadati, Yasaman, et al.
Published: (2024)
by: Saadati, Yasaman, et al.
Published: (2024)
TokenCake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications
by: Bian, Zhuohang, et al.
Published: (2025)
by: Bian, Zhuohang, et al.
Published: (2025)
FLEdge: Benchmarking Federated Machine Learning Applications in Edge Computing Systems
by: Woisetschläger, Herbert, et al.
Published: (2023)
by: Woisetschläger, Herbert, et al.
Published: (2023)
SPEAR++: Scaling Gradient Inversion via Sparsely-Used Dictionary Learning
by: Bakarsky, Alexander, et al.
Published: (2025)
by: Bakarsky, Alexander, et al.
Published: (2025)
Mobile Traffic Prediction at the Edge Through Distributed and Deep Transfer Learning
by: Petrella, Alfredo, et al.
Published: (2023)
by: Petrella, Alfredo, et al.
Published: (2023)
Accelerating Geo-distributed Machine Learning with Network-Aware Adaptive Tree and Auxiliary Route
by: Li, Zonghang, et al.
Published: (2024)
by: Li, Zonghang, et al.
Published: (2024)
A Taxonomy and Resolution Strategy for Client-Level Disagreements in Federated Learning
by: Rosendal, Daan, et al.
Published: (2026)
by: Rosendal, Daan, et al.
Published: (2026)
SI-ChainFL: Shapley-Incentivized Secure Federated Learning for High-Speed Rail Data Sharing
by: Zhao, Mingjie, et al.
Published: (2026)
by: Zhao, Mingjie, et al.
Published: (2026)
Uncertainty Estimation in Multi-Agent Distributed Learning for AI-Enabled Edge Devices
by: Radchenko, Gleb, et al.
Published: (2024)
by: Radchenko, Gleb, et al.
Published: (2024)
Bridging Generalization Gap of Heterogeneous Federated Clients Using Generative Models
by: Niu, Ziru, et al.
Published: (2025)
by: Niu, Ziru, et al.
Published: (2025)
A Survey on Efficient Federated Learning Methods for Foundation Model Training
by: Woisetschläger, Herbert, et al.
Published: (2024)
by: Woisetschläger, Herbert, et al.
Published: (2024)
Resource Heterogeneity-Aware and Utilization-Enhanced Scheduling for Deep Learning Clusters
by: Sultana, Abeda, et al.
Published: (2025)
by: Sultana, Abeda, et al.
Published: (2025)
Federated Learning for Traffic Flow Prediction with Synthetic Data Augmentation
by: Orozco, Fermin, et al.
Published: (2024)
by: Orozco, Fermin, et al.
Published: (2024)
ADF-LoRA: Alternating Low-Rank Aggregation for Decentralized Federated Fine-Tuning
by: Wang, Xiaoyu, et al.
Published: (2025)
by: Wang, Xiaoyu, et al.
Published: (2025)
A Full Compression Pipeline for Green Federated Learning in Communication-Constrained Environments
by: Colybes, Elouan, et al.
Published: (2026)
by: Colybes, Elouan, et al.
Published: (2026)
Federated Few-Shot Learning on Neuromorphic Hardware: An Empirical Study Across Physical Edge Nodes
by: Motta, Steven, et al.
Published: (2026)
by: Motta, Steven, et al.
Published: (2026)
Federated Fine-Tuning of LLMs on the Very Edge: The Good, the Bad, the Ugly
by: Woisetschläger, Herbert, et al.
Published: (2023)
by: Woisetschläger, Herbert, et al.
Published: (2023)
Similar Items
-
AutoDDL: Automatic Distributed Deep Learning with Near-Optimal Bandwidth Cost
by: Chen, Jinfan, et al.
Published: (2023) -
Chimera: Efficiently Training Large-Scale Neural Networks with Bidirectional Pipelines
by: Li, Shigang, et al.
Published: (2021) -
FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor Cores
by: Shi, Jinliang, et al.
Published: (2024) -
Libra: Unleashing GPU Heterogeneity for High-Performance Sparse Matrix Multiplication
by: Shi, Jinliang, et al.
Published: (2025) -
SparkAttention: High-Performance Multi-Head Attention for Large Models on Volta GPU Architecture
by: Xu, Youxuan, et al.
Published: (2025)