Convergence Analysis of SGD under Expected Smoothness
Fuente:
arXiv
Guardado en:
| Autores principales: | Kawamoto, Yuta, Iiduka, Hideaki |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Muon Converges under Heavy-Tailed Noise: Nonconvex Hölder-Smooth Empirical Risk Minimization
por: Iiduka, Hideaki
Publicado: (2026)
por: Iiduka, Hideaki
Publicado: (2026)
Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
por: Umeda, Hikaru, et al.
Publicado: (2025)
por: Umeda, Hikaru, et al.
Publicado: (2025)
Using Stochastic Gradient Descent to Smooth Nonconvex Functions: Analysis of Implicit Graduated Optimization
por: Sato, Naoki, et al.
Publicado: (2023)
por: Sato, Naoki, et al.
Publicado: (2023)
Increasing Batch Size Improves Convergence of Stochastic Gradient Descent with Momentum
por: Kamo, Keisuke, et al.
Publicado: (2025)
por: Kamo, Keisuke, et al.
Publicado: (2025)
Improved Convergence Rates of Muon Optimizer for Nonconvex Optimization
por: Nagashima, Shuntaro, et al.
Publicado: (2026)
por: Nagashima, Shuntaro, et al.
Publicado: (2026)
Faster Convergence of Riemannian Stochastic Gradient Descent with Increasing Batch Size
por: Oowada, Kanata, et al.
Publicado: (2025)
por: Oowada, Kanata, et al.
Publicado: (2025)
Both Asymptotic and Non-Asymptotic Convergence of Quasi-Hyperbolic Momentum using Increasing Batch Size
por: Imaizumi, Kento, et al.
Publicado: (2025)
por: Imaizumi, Kento, et al.
Publicado: (2025)
Convergence of Sharpness-Aware Minimization Algorithms using Increasing Batch Size and Decaying Learning Rate
por: Harada, Hinata, et al.
Publicado: (2024)
por: Harada, Hinata, et al.
Publicado: (2024)
Convergence Bound and Critical Batch Size of Muon Optimizer
por: Sato, Naoki, et al.
Publicado: (2025)
por: Sato, Naoki, et al.
Publicado: (2025)
Accelerating SGDM via Learning Rate and Batch Size Schedules: A Lyapunov-Based Analysis
por: Kondo, Yuichi, et al.
Publicado: (2025)
por: Kondo, Yuichi, et al.
Publicado: (2025)
Iteration and Stochastic First-order Oracle Complexities of Stochastic Gradient Descent using Constant and Decaying Learning Rates
por: Imaizumi, Kento, et al.
Publicado: (2024)
por: Imaizumi, Kento, et al.
Publicado: (2024)
Explicit and Implicit Graduated Optimization in Deep Neural Networks
por: Sato, Naoki, et al.
Publicado: (2024)
por: Sato, Naoki, et al.
Publicado: (2024)
Lipschitz Multiscale Deep Equilibrium Models: A Theoretically Guaranteed and Accelerated Approach
por: Sato, Naoki, et al.
Publicado: (2026)
por: Sato, Naoki, et al.
Publicado: (2026)
Adaptive Batch Size and Learning Rate Scheduler for Stochastic Gradient Descent Based on Minimization of Stochastic First-order Oracle Complexity
por: Umeda, Hikaru, et al.
Publicado: (2025)
por: Umeda, Hikaru, et al.
Publicado: (2025)
Relationship between Batch Size and Number of Steps Needed for Nonconvex Optimization of Stochastic Gradient Descent using Armijo Line Search
por: Tsukada, Yuki, et al.
Publicado: (2023)
por: Tsukada, Yuki, et al.
Publicado: (2023)
Increasing Both Batch Size and Learning Rate Accelerates Stochastic Gradient Descent
por: Umeda, Hikaru, et al.
Publicado: (2024)
por: Umeda, Hikaru, et al.
Publicado: (2024)
Momentum Does Not Reduce Stochastic Noise in Stochastic Gradient Descent
por: Sato, Naoki, et al.
Publicado: (2024)
por: Sato, Naoki, et al.
Publicado: (2024)
Scaled Conjugate Gradient Method for Nonconvex Optimization in Deep Neural Networks
por: Sato, Naoki, et al.
Publicado: (2024)
por: Sato, Naoki, et al.
Publicado: (2024)
Fast Last-Iterate Convergence of SGD in the Smooth Interpolation Regime
por: Attia, Amit, et al.
Publicado: (2025)
por: Attia, Amit, et al.
Publicado: (2025)
Convergent Privacy Loss of Noisy-SGD without Convexity and Smoothness
por: Chien, Eli, et al.
Publicado: (2024)
por: Chien, Eli, et al.
Publicado: (2024)
Diagonalisation SGD: Fast & Convergent SGD for Non-Differentiable Models via Reparameterisation and Smoothing
por: Wagner, Dominik, et al.
Publicado: (2024)
por: Wagner, Dominik, et al.
Publicado: (2024)
Convergence of Clipped-SGD for Convex $(L_0,L_1)$-Smooth Optimization with Heavy-Tailed Noise
por: Chezhegov, Savelii, et al.
Publicado: (2025)
por: Chezhegov, Savelii, et al.
Publicado: (2025)
Convergence Rates of Constrained Expected Improvement
por: Wang, Haowei, et al.
Publicado: (2025)
por: Wang, Haowei, et al.
Publicado: (2025)
PCDP-SGD: Improving the Convergence of Differentially Private SGD via Projection in Advance
por: Sha, Haichao, et al.
Publicado: (2023)
por: Sha, Haichao, et al.
Publicado: (2023)
Enhancing SignSGD: Small-Batch Convergence Analysis and a Hybrid Switching Strategy
por: Chen, Haoran, et al.
Publicado: (2026)
por: Chen, Haoran, et al.
Publicado: (2026)
MGDA Converges under Generalized Smoothness, Provably
por: Zhang, Qi, et al.
Publicado: (2024)
por: Zhang, Qi, et al.
Publicado: (2024)
Smoothed SGD for quantiles: Bahadur representation and Gaussian approximation
por: Chen, Likai, et al.
Publicado: (2025)
por: Chen, Likai, et al.
Publicado: (2025)
Bilevel Optimization under Unbounded Smoothness: A New Algorithm and Convergence Analysis
por: Hao, Jie, et al.
Publicado: (2024)
por: Hao, Jie, et al.
Publicado: (2024)
An Improved Privacy and Utility Analysis of Differentially Private SGD with Bounded Domain and Smooth Losses
por: Liang, Hao, et al.
Publicado: (2025)
por: Liang, Hao, et al.
Publicado: (2025)
On the Convergence of DP-SGD with Adaptive Clipping
por: Shulgin, Egor, et al.
Publicado: (2024)
por: Shulgin, Egor, et al.
Publicado: (2024)
Discovering Learning-Friendly Generation Orders for Sequential Computation
por: Sato, Yuta, et al.
Publicado: (2025)
por: Sato, Yuta, et al.
Publicado: (2025)
Faster Convergence of Local SGD for Over-Parameterized Models
por: Qin, Tiancheng, et al.
Publicado: (2022)
por: Qin, Tiancheng, et al.
Publicado: (2022)
Global Convergence of SGD On Two Layer Neural Nets
por: Gopalani, Pulkit, et al.
Publicado: (2022)
por: Gopalani, Pulkit, et al.
Publicado: (2022)
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees
por: Xie, Shengping, et al.
Publicado: (2025)
por: Xie, Shengping, et al.
Publicado: (2025)
High-Probability Convergence Guarantees of Decentralized SGD
por: Armacki, Aleksandar, et al.
Publicado: (2025)
por: Armacki, Aleksandar, et al.
Publicado: (2025)
Convergence Analysis of Randomized Subspace Normalized SGD under Heavy-Tailed Noise
por: Omiya, Gaku, et al.
Publicado: (2026)
por: Omiya, Gaku, et al.
Publicado: (2026)
Convergence of Steepest Descent and Adam under Non-Uniform Smoothness
por: Vaswani, Sharan, et al.
Publicado: (2026)
por: Vaswani, Sharan, et al.
Publicado: (2026)
Noise is All You Need: Private Second-Order Convergence of Noisy SGD
por: Avdiukhin, Dmitrii, et al.
Publicado: (2024)
por: Avdiukhin, Dmitrii, et al.
Publicado: (2024)
Convergence, Sticking and Escape: Stochastic Dynamics Near Critical Points in SGD
por: Dudukalov, Dmitry, et al.
Publicado: (2025)
por: Dudukalov, Dmitry, et al.
Publicado: (2025)
Mini-Batch Stochastic Halpern Algorithm for Nonexpansive Fixed point Problems
por: Iiduka, Hideaki
Publicado: (2026)
por: Iiduka, Hideaki
Publicado: (2026)
Ejemplares similares
-
Muon Converges under Heavy-Tailed Noise: Nonconvex Hölder-Smooth Empirical Risk Minimization
por: Iiduka, Hideaki
Publicado: (2026) -
Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
por: Umeda, Hikaru, et al.
Publicado: (2025) -
Using Stochastic Gradient Descent to Smooth Nonconvex Functions: Analysis of Implicit Graduated Optimization
por: Sato, Naoki, et al.
Publicado: (2023) -
Increasing Batch Size Improves Convergence of Stochastic Gradient Descent with Momentum
por: Kamo, Keisuke, et al.
Publicado: (2025) -
Improved Convergence Rates of Muon Optimizer for Nonconvex Optimization
por: Nagashima, Shuntaro, et al.
Publicado: (2026)