A Dynamic Weighting Strategy to Mitigate Worker Node Failure in Distributed Deep Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Yuesheng, Carr, Arielle |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dynamic Byzantine-Robust Learning: Adapting to Switching Byzantine Workers
von: Dorfman, Ron, et al.
Veröffentlicht: (2024)
von: Dorfman, Ron, et al.
Veröffentlicht: (2024)
Virtual Nodes Can Help: Tackling Distribution Shifts in Federated Graph Learning
von: Fu, Xingbo, et al.
Veröffentlicht: (2024)
von: Fu, Xingbo, et al.
Veröffentlicht: (2024)
Communication-Efficient Distributed Deep Learning via Federated Dynamic Averaging
von: Theologitis, Michail, et al.
Veröffentlicht: (2024)
von: Theologitis, Michail, et al.
Veröffentlicht: (2024)
Deal: Distributed End-to-End GNN Inference for All Nodes
von: Chen, Shiyang, et al.
Veröffentlicht: (2025)
von: Chen, Shiyang, et al.
Veröffentlicht: (2025)
AntDT: A Self-Adaptive Distributed Training Framework for Leader and Straggler Nodes
von: Xiao, Youshao, et al.
Veröffentlicht: (2024)
von: Xiao, Youshao, et al.
Veröffentlicht: (2024)
Going Forward-Forward in Distributed Deep Learning
von: Aktemur, Ege, et al.
Veröffentlicht: (2024)
von: Aktemur, Ege, et al.
Veröffentlicht: (2024)
Distributed Deep Learning using Stochastic Gradient Staleness
von: Pham, Viet Hoang, et al.
Veröffentlicht: (2025)
von: Pham, Viet Hoang, et al.
Veröffentlicht: (2025)
Improved Quantization Strategies for Managing Heavy-tailed Gradients in Distributed Learning
von: Yan, Guangfeng, et al.
Veröffentlicht: (2024)
von: Yan, Guangfeng, et al.
Veröffentlicht: (2024)
Optimizing the Optimal Weighted Average: Efficient Distributed Sparse Classification
von: Lu, Fred, et al.
Veröffentlicht: (2024)
von: Lu, Fred, et al.
Veröffentlicht: (2024)
NEST: Network- and Memory-Aware Device Placement For Distributed Deep Learning
von: Wang, Irene, et al.
Veröffentlicht: (2026)
von: Wang, Irene, et al.
Veröffentlicht: (2026)
SparDL: Distributed Deep Learning Training with Efficient Sparse Communication
von: Zhao, Minjun, et al.
Veröffentlicht: (2023)
von: Zhao, Minjun, et al.
Veröffentlicht: (2023)
Empowering Federated Learning with Implicit Gossiping: Mitigating Connection Unreliability Amidst Unknown and Arbitrary Dynamics
von: Xiang, Ming, et al.
Veröffentlicht: (2024)
von: Xiang, Ming, et al.
Veröffentlicht: (2024)
FedRFQ: Prototype-Based Federated Learning with Reduced Redundancy, Minimal Failure, and Enhanced Quality
von: Yan, Biwei, et al.
Veröffentlicht: (2024)
von: Yan, Biwei, et al.
Veröffentlicht: (2024)
GRAWA: Gradient-based Weighted Averaging for Distributed Training of Deep Learning Models
von: Dimlioglu, Tolga, et al.
Veröffentlicht: (2024)
von: Dimlioglu, Tolga, et al.
Veröffentlicht: (2024)
Preserving Near-Optimal Gradient Sparsification Cost for Scalable Distributed Deep Learning
von: Yoon, Daegun, et al.
Veröffentlicht: (2024)
von: Yoon, Daegun, et al.
Veröffentlicht: (2024)
Communication-Efficient Distributed Training for Collaborative Flat Optima Recovery in Deep Learning
von: Dimlioglu, Tolga, et al.
Veröffentlicht: (2025)
von: Dimlioglu, Tolga, et al.
Veröffentlicht: (2025)
Prediction-Assisted Online Distributed Deep Learning Workload Scheduling in GPU Clusters
von: Luo, Ziyue, et al.
Veröffentlicht: (2025)
von: Luo, Ziyue, et al.
Veröffentlicht: (2025)
Federated Learning Optimization: A Comparative Study of Data and Model Exchange Strategies in Dynamic Networks
von: Luqman, Alka, et al.
Veröffentlicht: (2024)
von: Luqman, Alka, et al.
Veröffentlicht: (2024)
Time-Series Learning for Proactive Fault Prediction in Distributed Systems with Deep Neural Structures
von: Wang, Yang, et al.
Veröffentlicht: (2025)
von: Wang, Yang, et al.
Veröffentlicht: (2025)
FedClust: Optimizing Federated Learning on Non-IID Data through Weight-Driven Client Clustering
von: Islam, Md Sirajul, et al.
Veröffentlicht: (2024)
von: Islam, Md Sirajul, et al.
Veröffentlicht: (2024)
Communication-Efficient Personalized Distributed Learning with Data and Node Heterogeneity
von: Tian, Zhuojun, et al.
Veröffentlicht: (2025)
von: Tian, Zhuojun, et al.
Veröffentlicht: (2025)
Aryl: An Elastic Cluster Scheduler for Deep Learning
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
An Experimental Comparison of Partitioning Strategies for Distributed Graph Neural Network Training
von: Merkel, Nikolai, et al.
Veröffentlicht: (2023)
von: Merkel, Nikolai, et al.
Veröffentlicht: (2023)
First-Order Softmax Weighted Switching Gradient Method for Distributed Stochastic Minimax Optimization with Stochastic Constraints
von: Luo, Zhankun, et al.
Veröffentlicht: (2026)
von: Luo, Zhankun, et al.
Veröffentlicht: (2026)
CAFE: Carbon-Aware Federated Learning in Geographically Distributed Data Centers
von: Bian, Jieming, et al.
Veröffentlicht: (2023)
von: Bian, Jieming, et al.
Veröffentlicht: (2023)
Traversal Learning: A Lossless And Efficient Distributed Learning Framework
von: Batbaatar, Erdenebileg, et al.
Veröffentlicht: (2025)
von: Batbaatar, Erdenebileg, et al.
Veröffentlicht: (2025)
Reinforcement Learning-based Adaptive Mitigation of Uncorrected DRAM Errors in the Field
von: Boixaderas, Isaac, et al.
Veröffentlicht: (2024)
von: Boixaderas, Isaac, et al.
Veröffentlicht: (2024)
Bandwidth-Aware and Overlap-Weighted Compression for Communication-Efficient Federated Learning
von: Tang, Zichen, et al.
Veröffentlicht: (2024)
von: Tang, Zichen, et al.
Veröffentlicht: (2024)
Communication-Efficient Distributed Learning with Local Immediate Error Compensation
von: Cheng, Yifei, et al.
Veröffentlicht: (2024)
von: Cheng, Yifei, et al.
Veröffentlicht: (2024)
Mask-Encoded Sparsification: Mitigating Biased Gradients in Communication-Efficient Split Learning
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
Robust Federated Learning Mitigates Client-side Training Data Distribution Inference Attacks
von: Xu, Yichang, et al.
Veröffentlicht: (2024)
von: Xu, Yichang, et al.
Veröffentlicht: (2024)
Loss- and Reward-Weighting for Efficient Distributed Reinforcement Learning
von: Holen, Martin, et al.
Veröffentlicht: (2023)
von: Holen, Martin, et al.
Veröffentlicht: (2023)
Federated Learning Under Temporal Drift -- Mitigating Catastrophic Forgetting via Experience Replay
von: Kokkula, Sahasra, et al.
Veröffentlicht: (2026)
von: Kokkula, Sahasra, et al.
Veröffentlicht: (2026)
Communication-Efficient Federated Learning through Adaptive Weight Clustering and Server-Side Distillation
von: Tsouvalas, Vasileios, et al.
Veröffentlicht: (2024)
von: Tsouvalas, Vasileios, et al.
Veröffentlicht: (2024)
Incentivizing Permissionless Distributed Learning of LLMs
von: Lidin, Joel, et al.
Veröffentlicht: (2025)
von: Lidin, Joel, et al.
Veröffentlicht: (2025)
Tackling Resource-Constrained and Data-Heterogeneity in Federated Learning with Double-Weight Sparse Pack
von: Yang, Qiantao, et al.
Veröffentlicht: (2026)
von: Yang, Qiantao, et al.
Veröffentlicht: (2026)
DA-PFL: Dynamic Affinity Aggregation for Personalized Federated Learning
von: Yang, Xu, et al.
Veröffentlicht: (2024)
von: Yang, Xu, et al.
Veröffentlicht: (2024)
GraNNDis: Efficient Unified Distributed Training Framework for Deep GNNs on Large Clusters
von: Song, Jaeyong, et al.
Veröffentlicht: (2023)
von: Song, Jaeyong, et al.
Veröffentlicht: (2023)
Mitigating Catastrophic Forgetting with Adaptive Transformer Block Expansion in Federated Fine-Tuning
von: Huo, Yujia, et al.
Veröffentlicht: (2025)
von: Huo, Yujia, et al.
Veröffentlicht: (2025)
ECHO-2: A Large-Scale Distributed Rollout Framework for Cost-Efficient Reinforcement Learning
von: Song, Jingwei, et al.
Veröffentlicht: (2026)
von: Song, Jingwei, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Dynamic Byzantine-Robust Learning: Adapting to Switching Byzantine Workers
von: Dorfman, Ron, et al.
Veröffentlicht: (2024) -
Virtual Nodes Can Help: Tackling Distribution Shifts in Federated Graph Learning
von: Fu, Xingbo, et al.
Veröffentlicht: (2024) -
Communication-Efficient Distributed Deep Learning via Federated Dynamic Averaging
von: Theologitis, Michail, et al.
Veröffentlicht: (2024) -
Deal: Distributed End-to-End GNN Inference for All Nodes
von: Chen, Shiyang, et al.
Veröffentlicht: (2025) -
AntDT: A Self-Adaptive Distributed Training Framework for Leader and Straggler Nodes
von: Xiao, Youshao, et al.
Veröffentlicht: (2024)