Distributed Deep Learning using Stochastic Gradient Staleness
Fuente:
arXiv
Saved in:
| Main Authors: | Pham, Viet Hoang, Ahn, Hyo-Sung |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Distributed Stochastic Gradient Descent with Staleness: A Stochastic Delay Differential Equation Based Framework
by: Yu, Siyuan, et al.
Published: (2024)
by: Yu, Siyuan, et al.
Published: (2024)
FedStaleWeight: Buffered Asynchronous Federated Learning with Fair Aggregation via Staleness Reweighting
by: Ma, Jeffrey, et al.
Published: (2024)
by: Ma, Jeffrey, et al.
Published: (2024)
Tackling Intertwined Data and Device Heterogeneities in Federated Learning with Unlimited Staleness
by: Wang, Haoming, et al.
Published: (2023)
by: Wang, Haoming, et al.
Published: (2023)
MSPipe: Efficient Temporal GNN Training via Staleness-Aware Pipeline
by: Sheng, Guangming, et al.
Published: (2024)
by: Sheng, Guangming, et al.
Published: (2024)
Preserving Near-Optimal Gradient Sparsification Cost for Scalable Distributed Deep Learning
by: Yoon, Daegun, et al.
Published: (2024)
by: Yoon, Daegun, et al.
Published: (2024)
First-Order Softmax Weighted Switching Gradient Method for Distributed Stochastic Minimax Optimization with Stochastic Constraints
by: Luo, Zhankun, et al.
Published: (2026)
by: Luo, Zhankun, et al.
Published: (2026)
ToFU: Transforming How Federated Learning Systems Forget User Data
by: Tran, Van-Tuan, et al.
Published: (2025)
by: Tran, Van-Tuan, et al.
Published: (2025)
Trustworthiness of Stochastic Gradient Descent in Distributed Learning
by: Li, Hongyang, et al.
Published: (2024)
by: Li, Hongyang, et al.
Published: (2024)
Improved Quantization Strategies for Managing Heavy-tailed Gradients in Distributed Learning
by: Yan, Guangfeng, et al.
Published: (2024)
by: Yan, Guangfeng, et al.
Published: (2024)
Distributed Learning based on 1-Bit Gradient Coding in the Presence of Stragglers
by: Li, Chengxi, et al.
Published: (2024)
by: Li, Chengxi, et al.
Published: (2024)
ADP-VRSGP: Decentralized Learning with Adaptive Differential Privacy via Variance-Reduced Stochastic Gradient Push
by: Wu, Xiaoming, et al.
Published: (2025)
by: Wu, Xiaoming, et al.
Published: (2025)
Going Forward-Forward in Distributed Deep Learning
by: Aktemur, Ege, et al.
Published: (2024)
by: Aktemur, Ege, et al.
Published: (2024)
Mitigating Staleness in Asynchronous Pipeline Parallelism via Basis Rotation
by: Jung, Hyunji, et al.
Published: (2026)
by: Jung, Hyunji, et al.
Published: (2026)
NEST: Network- and Memory-Aware Device Placement For Distributed Deep Learning
by: Wang, Irene, et al.
Published: (2026)
by: Wang, Irene, et al.
Published: (2026)
Communication-Efficient Distributed Deep Learning via Federated Dynamic Averaging
by: Theologitis, Michail, et al.
Published: (2024)
by: Theologitis, Michail, et al.
Published: (2024)
SparDL: Distributed Deep Learning Training with Efficient Sparse Communication
by: Zhao, Minjun, et al.
Published: (2023)
by: Zhao, Minjun, et al.
Published: (2023)
Effectiveness of Distributed Gradient Descent with Local Steps for Overparameterized Models
by: Zhu, Heng, et al.
Published: (2024)
by: Zhu, Heng, et al.
Published: (2024)
Asynch-SGBDT: Asynchronous Parallel Stochastic Gradient Boosting Decision Tree based on Parameters Server
by: Daning, Cheng, et al.
Published: (2018)
by: Daning, Cheng, et al.
Published: (2018)
Communication-Efficient Distributed Training for Collaborative Flat Optima Recovery in Deep Learning
by: Dimlioglu, Tolga, et al.
Published: (2025)
by: Dimlioglu, Tolga, et al.
Published: (2025)
Prediction-Assisted Online Distributed Deep Learning Workload Scheduling in GPU Clusters
by: Luo, Ziyue, et al.
Published: (2025)
by: Luo, Ziyue, et al.
Published: (2025)
GRAWA: Gradient-based Weighted Averaging for Distributed Training of Deep Learning Models
by: Dimlioglu, Tolga, et al.
Published: (2024)
by: Dimlioglu, Tolga, et al.
Published: (2024)
Time-Series Learning for Proactive Fault Prediction in Distributed Systems with Deep Neural Structures
by: Wang, Yang, et al.
Published: (2025)
by: Wang, Yang, et al.
Published: (2025)
A Dynamic Weighting Strategy to Mitigate Worker Node Failure in Distributed Deep Learning
by: Xu, Yuesheng, et al.
Published: (2024)
by: Xu, Yuesheng, et al.
Published: (2024)
High-Dimensional Sparse Data Low-rank Representation via Accelerated Asynchronous Parallel Stochastic Gradient Descent
by: Hu, Qicong, et al.
Published: (2024)
by: Hu, Qicong, et al.
Published: (2024)
Federated Learning on Stochastic Neural Networks
by: Tang, Jingqiao, et al.
Published: (2025)
by: Tang, Jingqiao, et al.
Published: (2025)
MiCRO: Near-Zero Cost Gradient Sparsification for Scaling and Accelerating Distributed DNN Training
by: Yoon, Daegun, et al.
Published: (2023)
by: Yoon, Daegun, et al.
Published: (2023)
Maya: Optimizing Deep Learning Training Workloads using GPU Runtime Emulation
by: Yarlagadda, Srihas, et al.
Published: (2025)
by: Yarlagadda, Srihas, et al.
Published: (2025)
Robust Decentralized Learning with Local Updates and Gradient Tracking
by: Ghiasvand, Sajjad, et al.
Published: (2024)
by: Ghiasvand, Sajjad, et al.
Published: (2024)
FedAgg: Adaptive Federated Learning with Aggregated Gradients
by: Yuan, Wenhao, et al.
Published: (2023)
by: Yuan, Wenhao, et al.
Published: (2023)
Buffer-based Gradient Projection for Continual Federated Learning
by: Dai, Shenghong, et al.
Published: (2024)
by: Dai, Shenghong, et al.
Published: (2024)
Cornserve: A Distributed Serving System for Any-to-Any Multimodal Models
by: Chung, Jae-Won, et al.
Published: (2026)
by: Chung, Jae-Won, et al.
Published: (2026)
A-3PO: Accelerating Asynchronous LLM Training with Staleness-aware Proximal Policy Approximation
by: Li, Xiaocan, et al.
Published: (2025)
by: Li, Xiaocan, et al.
Published: (2025)
On the Convergence of Continual Federated Learning Using Incrementally Aggregated Gradients
by: Keshri, Satish Kumar, et al.
Published: (2024)
by: Keshri, Satish Kumar, et al.
Published: (2024)
Accelerating Federated Learning by Selecting Beneficial Herd of Local Gradients
by: Luo, Ping, et al.
Published: (2024)
by: Luo, Ping, et al.
Published: (2024)
An Efficient Gradient-Aware Error-Bounded Lossy Compressor for Federated Learning
by: Ye, Zhijing, et al.
Published: (2025)
by: Ye, Zhijing, et al.
Published: (2025)
Age-of-Gradient Updates for Federated Learning over Random Access Channels
by: Wu, Yu Heng, et al.
Published: (2024)
by: Wu, Yu Heng, et al.
Published: (2024)
Local Gradient Regulation Stabilizes Federated Learning under Client Heterogeneity
by: Luo, Ping, et al.
Published: (2026)
by: Luo, Ping, et al.
Published: (2026)
Deep Reinforcement Learning for System-on-Chip: Myths and Realities
by: Sung, Tegg Taekyong, et al.
Published: (2022)
by: Sung, Tegg Taekyong, et al.
Published: (2022)
FedQS: Optimizing Gradient and Model Aggregation for Semi-Asynchronous Federated Learning
by: Li, Yunbo, et al.
Published: (2025)
by: Li, Yunbo, et al.
Published: (2025)
Mask-Encoded Sparsification: Mitigating Biased Gradients in Communication-Efficient Split Learning
by: Zhou, Wenxuan, et al.
Published: (2024)
by: Zhou, Wenxuan, et al.
Published: (2024)
Similar Items
-
Distributed Stochastic Gradient Descent with Staleness: A Stochastic Delay Differential Equation Based Framework
by: Yu, Siyuan, et al.
Published: (2024) -
FedStaleWeight: Buffered Asynchronous Federated Learning with Fair Aggregation via Staleness Reweighting
by: Ma, Jeffrey, et al.
Published: (2024) -
Tackling Intertwined Data and Device Heterogeneities in Federated Learning with Unlimited Staleness
by: Wang, Haoming, et al.
Published: (2023) -
MSPipe: Efficient Temporal GNN Training via Staleness-Aware Pipeline
by: Sheng, Guangming, et al.
Published: (2024) -
Preserving Near-Optimal Gradient Sparsification Cost for Scalable Distributed Deep Learning
by: Yoon, Daegun, et al.
Published: (2024)