A Survey of Optimization Methods for Training DL Models: Theoretical Perspective on Convergence and Generalization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Jing, Choromanska, Anna |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GRAWA: Gradient-based Weighted Averaging for Distributed Training of Deep Learning Models
von: Dimlioglu, Tolga, et al.
Veröffentlicht: (2024)
von: Dimlioglu, Tolga, et al.
Veröffentlicht: (2024)
LoCoDL: Communication-Efficient Distributed Learning with Local Training and Compression
von: Condat, Laurent, et al.
Veröffentlicht: (2024)
von: Condat, Laurent, et al.
Veröffentlicht: (2024)
A Double Tracking Method for Optimization with Decentralized Generalized Orthogonality Constraints
von: Wang, Lei, et al.
Veröffentlicht: (2024)
von: Wang, Lei, et al.
Veröffentlicht: (2024)
A New Theoretical Perspective on Data Heterogeneity in Federated Optimization
von: Wang, Jiayi, et al.
Veröffentlicht: (2024)
von: Wang, Jiayi, et al.
Veröffentlicht: (2024)
Towards a Better Theoretical Understanding of Independent Subnetwork Training
von: Shulgin, Egor, et al.
Veröffentlicht: (2023)
von: Shulgin, Egor, et al.
Veröffentlicht: (2023)
Convergence of Sign-based Random Reshuffling Algorithms for Nonconvex Optimization
von: Qin, Zhen, et al.
Veröffentlicht: (2023)
von: Qin, Zhen, et al.
Veröffentlicht: (2023)
Achieving Near-Optimal Convergence for Distributed Minimax Optimization with Adaptive Stepsizes
von: Huang, Yan, et al.
Veröffentlicht: (2024)
von: Huang, Yan, et al.
Veröffentlicht: (2024)
Accelerating Distributed Optimization: A Primal-Dual Perspective on Local Steps
von: Yang, Junchi, et al.
Veröffentlicht: (2024)
von: Yang, Junchi, et al.
Veröffentlicht: (2024)
On Principled Local Optimization Methods for Federated Learning
von: Yuan, Honglin
Veröffentlicht: (2024)
von: Yuan, Honglin
Veröffentlicht: (2024)
Activations and Gradients Compression for Model-Parallel Training
von: Rudakov, Mikhail, et al.
Veröffentlicht: (2024)
von: Rudakov, Mikhail, et al.
Veröffentlicht: (2024)
Accelerated Methods with Compressed Communications for Distributed Optimization Problems under Data Similarity
von: Bylinkin, Dmitry, et al.
Veröffentlicht: (2024)
von: Bylinkin, Dmitry, et al.
Veröffentlicht: (2024)
A Hybrid Stochastic Gradient Tracking Method for Distributed Online Optimization Over Time-Varying Directed Networks
von: Shi, Xinli, et al.
Veröffentlicht: (2025)
von: Shi, Xinli, et al.
Veröffentlicht: (2025)
FIARSE: Model-Heterogeneous Federated Learning via Importance-Aware Submodel Extraction
von: Wu, Feijie, et al.
Veröffentlicht: (2024)
von: Wu, Feijie, et al.
Veröffentlicht: (2024)
Communication-Efficient Distributed Training for Collaborative Flat Optima Recovery in Deep Learning
von: Dimlioglu, Tolga, et al.
Veröffentlicht: (2025)
von: Dimlioglu, Tolga, et al.
Veröffentlicht: (2025)
S$^3$LDBO: A Snapshot Single-Loop Algorithm for Decentralized Bilevel Optimization
von: Yin, Chao, et al.
Veröffentlicht: (2026)
von: Yin, Chao, et al.
Veröffentlicht: (2026)
Communication-Efficient Federated Optimization over Semi-Decentralized Networks
von: Wang, He, et al.
Veröffentlicht: (2023)
von: Wang, He, et al.
Veröffentlicht: (2023)
A Penalty-Based Method for Communication-Efficient Decentralized Bilevel Programming
von: Nazari, Parvin, et al.
Veröffentlicht: (2022)
von: Nazari, Parvin, et al.
Veröffentlicht: (2022)
A Single-Loop Algorithm for Decentralized Bilevel Optimization
von: Dong, Youran, et al.
Veröffentlicht: (2023)
von: Dong, Youran, et al.
Veröffentlicht: (2023)
Local Methods with Adaptivity via Scaling
von: Chezhegov, Savelii, et al.
Veröffentlicht: (2024)
von: Chezhegov, Savelii, et al.
Veröffentlicht: (2024)
Birch SGD: A Tree Graph Framework for Local and Asynchronous SGD Methods
von: Tyurin, Alexander, et al.
Veröffentlicht: (2025)
von: Tyurin, Alexander, et al.
Veröffentlicht: (2025)
CEDAS: A Compressed Decentralized Stochastic Gradient Method with Improved Convergence
von: Huang, Kun, et al.
Veröffentlicht: (2023)
von: Huang, Kun, et al.
Veröffentlicht: (2023)
A Stochastic Approximation Approach for Efficient Decentralized Optimization on Random Networks
von: Yau, Chung-Yiu, et al.
Veröffentlicht: (2024)
von: Yau, Chung-Yiu, et al.
Veröffentlicht: (2024)
Efficient Adaptive Federated Optimization
von: Lee, Su Hyeong, et al.
Veröffentlicht: (2024)
von: Lee, Su Hyeong, et al.
Veröffentlicht: (2024)
CONGO: Compressive Online Gradient Optimization
von: Carleton, Jeremy, et al.
Veröffentlicht: (2024)
von: Carleton, Jeremy, et al.
Veröffentlicht: (2024)
Ringmaster LMO: Asynchronous Linear Minimization Oracle Momentum Method
von: Sadiev, Abdurakhmon, et al.
Veröffentlicht: (2026)
von: Sadiev, Abdurakhmon, et al.
Veröffentlicht: (2026)
Optimizing Stochastic Gradient Push under Broadcast Communications
von: Nguyen, Tuan, et al.
Veröffentlicht: (2026)
von: Nguyen, Tuan, et al.
Veröffentlicht: (2026)
GradSkip: Communication-Accelerated Local Gradient Methods with Better Computational Complexity
von: Maranjyan, Artavazd, et al.
Veröffentlicht: (2022)
von: Maranjyan, Artavazd, et al.
Veröffentlicht: (2022)
Lower Bounds and Accelerated Algorithms in Distributed Stochastic Optimization with Communication Compression
von: He, Yutong, et al.
Veröffentlicht: (2023)
von: He, Yutong, et al.
Veröffentlicht: (2023)
Unbiased Compression Saves Communication in Distributed Optimization: When and How Much?
von: He, Yutong, et al.
Veröffentlicht: (2023)
von: He, Yutong, et al.
Veröffentlicht: (2023)
Communication-Efficient Federated Bilevel Optimization with Local and Global Lower Level Problems
von: Li, Junyi, et al.
Veröffentlicht: (2023)
von: Li, Junyi, et al.
Veröffentlicht: (2023)
Rescaled Asynchronous SGD: Optimal Distributed Optimization under Data and System Heterogeneity
von: Mahran, Ammar, et al.
Veröffentlicht: (2026)
von: Mahran, Ammar, et al.
Veröffentlicht: (2026)
AGD: an Auto-switchable Optimizer using Stepwise Gradient Difference for Preconditioning Matrix
von: Yue, Yun, et al.
Veröffentlicht: (2023)
von: Yue, Yun, et al.
Veröffentlicht: (2023)
Proving the Limited Scalability of Centralized Distributed Optimization via a New Lower Bound Construction
von: Tyurin, Alexander
Veröffentlicht: (2025)
von: Tyurin, Alexander
Veröffentlicht: (2025)
Rennala MVR: Improved Time Complexity for Parallel Stochastic Optimization via Momentum-Based Variance Reduction
von: Tovmasyan, Zhirayr, et al.
Veröffentlicht: (2026)
von: Tovmasyan, Zhirayr, et al.
Veröffentlicht: (2026)
Provable Model-Parallel Distributed Principal Component Analysis with Parallel Deflation
von: Liao, Fangshuo, et al.
Veröffentlicht: (2025)
von: Liao, Fangshuo, et al.
Veröffentlicht: (2025)
LOSCAR-SGD: Local SGD with Communication-Computation Overlap and Delay-Corrected Sparse Model Averaging
von: Maziane, Yassine, et al.
Veröffentlicht: (2026)
von: Maziane, Yassine, et al.
Veröffentlicht: (2026)
Towards Dynamic Resource Allocation and Client Scheduling in Hierarchical Federated Learning: A Two-Phase Deep Reinforcement Learning Approach
von: Chen, Xiaojing, et al.
Veröffentlicht: (2024)
von: Chen, Xiaojing, et al.
Veröffentlicht: (2024)
Demystifying Why Local Aggregation Helps: Convergence Analysis of Hierarchical SGD
von: Wang, Jiayi, et al.
Veröffentlicht: (2020)
von: Wang, Jiayi, et al.
Veröffentlicht: (2020)
Tailoring Gradient Methods for Differentially-Private Distributed Optimization
von: Wang, Yongqiang, et al.
Veröffentlicht: (2022)
von: Wang, Yongqiang, et al.
Veröffentlicht: (2022)
Decentralized Nonconvex Optimization under Heavy-Tailed Noise: Normalization and Optimal Convergence
von: Yu, Shuhua, et al.
Veröffentlicht: (2025)
von: Yu, Shuhua, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
GRAWA: Gradient-based Weighted Averaging for Distributed Training of Deep Learning Models
von: Dimlioglu, Tolga, et al.
Veröffentlicht: (2024) -
LoCoDL: Communication-Efficient Distributed Learning with Local Training and Compression
von: Condat, Laurent, et al.
Veröffentlicht: (2024) -
A Double Tracking Method for Optimization with Decentralized Generalized Orthogonality Constraints
von: Wang, Lei, et al.
Veröffentlicht: (2024) -
A New Theoretical Perspective on Data Heterogeneity in Federated Optimization
von: Wang, Jiayi, et al.
Veröffentlicht: (2024) -
Towards a Better Theoretical Understanding of Independent Subnetwork Training
von: Shulgin, Egor, et al.
Veröffentlicht: (2023)