Gespeichert in:
| Hauptverfasser: | Kawamoto, Yuta, Iiduka, Hideaki |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2510.20608 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Muon Converges under Heavy-Tailed Noise: Nonconvex Hölder-Smooth Empirical Risk Minimization
von: Iiduka, Hideaki
Veröffentlicht: (2026)
von: Iiduka, Hideaki
Veröffentlicht: (2026)
Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
von: Umeda, Hikaru, et al.
Veröffentlicht: (2025)
von: Umeda, Hikaru, et al.
Veröffentlicht: (2025)
Using Stochastic Gradient Descent to Smooth Nonconvex Functions: Analysis of Implicit Graduated Optimization
von: Sato, Naoki, et al.
Veröffentlicht: (2023)
von: Sato, Naoki, et al.
Veröffentlicht: (2023)
Increasing Batch Size Improves Convergence of Stochastic Gradient Descent with Momentum
von: Kamo, Keisuke, et al.
Veröffentlicht: (2025)
von: Kamo, Keisuke, et al.
Veröffentlicht: (2025)
Improved Convergence Rates of Muon Optimizer for Nonconvex Optimization
von: Nagashima, Shuntaro, et al.
Veröffentlicht: (2026)
von: Nagashima, Shuntaro, et al.
Veröffentlicht: (2026)
Faster Convergence of Riemannian Stochastic Gradient Descent with Increasing Batch Size
von: Oowada, Kanata, et al.
Veröffentlicht: (2025)
von: Oowada, Kanata, et al.
Veröffentlicht: (2025)
Both Asymptotic and Non-Asymptotic Convergence of Quasi-Hyperbolic Momentum using Increasing Batch Size
von: Imaizumi, Kento, et al.
Veröffentlicht: (2025)
von: Imaizumi, Kento, et al.
Veröffentlicht: (2025)
Convergence of Sharpness-Aware Minimization Algorithms using Increasing Batch Size and Decaying Learning Rate
von: Harada, Hinata, et al.
Veröffentlicht: (2024)
von: Harada, Hinata, et al.
Veröffentlicht: (2024)
Convergence Bound and Critical Batch Size of Muon Optimizer
von: Sato, Naoki, et al.
Veröffentlicht: (2025)
von: Sato, Naoki, et al.
Veröffentlicht: (2025)
Accelerating SGDM via Learning Rate and Batch Size Schedules: A Lyapunov-Based Analysis
von: Kondo, Yuichi, et al.
Veröffentlicht: (2025)
von: Kondo, Yuichi, et al.
Veröffentlicht: (2025)
Iteration and Stochastic First-order Oracle Complexities of Stochastic Gradient Descent using Constant and Decaying Learning Rates
von: Imaizumi, Kento, et al.
Veröffentlicht: (2024)
von: Imaizumi, Kento, et al.
Veröffentlicht: (2024)
Explicit and Implicit Graduated Optimization in Deep Neural Networks
von: Sato, Naoki, et al.
Veröffentlicht: (2024)
von: Sato, Naoki, et al.
Veröffentlicht: (2024)
Lipschitz Multiscale Deep Equilibrium Models: A Theoretically Guaranteed and Accelerated Approach
von: Sato, Naoki, et al.
Veröffentlicht: (2026)
von: Sato, Naoki, et al.
Veröffentlicht: (2026)
Adaptive Batch Size and Learning Rate Scheduler for Stochastic Gradient Descent Based on Minimization of Stochastic First-order Oracle Complexity
von: Umeda, Hikaru, et al.
Veröffentlicht: (2025)
von: Umeda, Hikaru, et al.
Veröffentlicht: (2025)
Relationship between Batch Size and Number of Steps Needed for Nonconvex Optimization of Stochastic Gradient Descent using Armijo Line Search
von: Tsukada, Yuki, et al.
Veröffentlicht: (2023)
von: Tsukada, Yuki, et al.
Veröffentlicht: (2023)
Increasing Both Batch Size and Learning Rate Accelerates Stochastic Gradient Descent
von: Umeda, Hikaru, et al.
Veröffentlicht: (2024)
von: Umeda, Hikaru, et al.
Veröffentlicht: (2024)
Momentum Does Not Reduce Stochastic Noise in Stochastic Gradient Descent
von: Sato, Naoki, et al.
Veröffentlicht: (2024)
von: Sato, Naoki, et al.
Veröffentlicht: (2024)
Scaled Conjugate Gradient Method for Nonconvex Optimization in Deep Neural Networks
von: Sato, Naoki, et al.
Veröffentlicht: (2024)
von: Sato, Naoki, et al.
Veröffentlicht: (2024)
Diagonalisation SGD: Fast & Convergent SGD for Non-Differentiable Models via Reparameterisation and Smoothing
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
Fast Last-Iterate Convergence of SGD in the Smooth Interpolation Regime
von: Attia, Amit, et al.
Veröffentlicht: (2025)
von: Attia, Amit, et al.
Veröffentlicht: (2025)
Convergent Privacy Loss of Noisy-SGD without Convexity and Smoothness
von: Chien, Eli, et al.
Veröffentlicht: (2024)
von: Chien, Eli, et al.
Veröffentlicht: (2024)
Convergence of Clipped-SGD for Convex $(L_0,L_1)$-Smooth Optimization with Heavy-Tailed Noise
von: Chezhegov, Savelii, et al.
Veröffentlicht: (2025)
von: Chezhegov, Savelii, et al.
Veröffentlicht: (2025)
Convergence Rates of Constrained Expected Improvement
von: Wang, Haowei, et al.
Veröffentlicht: (2025)
von: Wang, Haowei, et al.
Veröffentlicht: (2025)
Discovering Learning-Friendly Generation Orders for Sequential Computation
von: Sato, Yuta, et al.
Veröffentlicht: (2025)
von: Sato, Yuta, et al.
Veröffentlicht: (2025)
PCDP-SGD: Improving the Convergence of Differentially Private SGD via Projection in Advance
von: Sha, Haichao, et al.
Veröffentlicht: (2023)
von: Sha, Haichao, et al.
Veröffentlicht: (2023)
Smoothed SGD for quantiles: Bahadur representation and Gaussian approximation
von: Chen, Likai, et al.
Veröffentlicht: (2025)
von: Chen, Likai, et al.
Veröffentlicht: (2025)
MGDA Converges under Generalized Smoothness, Provably
von: Zhang, Qi, et al.
Veröffentlicht: (2024)
von: Zhang, Qi, et al.
Veröffentlicht: (2024)
On the Convergence of DP-SGD with Adaptive Clipping
von: Shulgin, Egor, et al.
Veröffentlicht: (2024)
von: Shulgin, Egor, et al.
Veröffentlicht: (2024)
Enhancing SignSGD: Small-Batch Convergence Analysis and a Hybrid Switching Strategy
von: Chen, Haoran, et al.
Veröffentlicht: (2026)
von: Chen, Haoran, et al.
Veröffentlicht: (2026)
An Improved Privacy and Utility Analysis of Differentially Private SGD with Bounded Domain and Smooth Losses
von: Liang, Hao, et al.
Veröffentlicht: (2025)
von: Liang, Hao, et al.
Veröffentlicht: (2025)
Bilevel Optimization under Unbounded Smoothness: A New Algorithm and Convergence Analysis
von: Hao, Jie, et al.
Veröffentlicht: (2024)
von: Hao, Jie, et al.
Veröffentlicht: (2024)
Mini-Batch Stochastic Halpern Algorithm for Nonexpansive Fixed point Problems
von: Iiduka, Hideaki
Veröffentlicht: (2026)
von: Iiduka, Hideaki
Veröffentlicht: (2026)
Mini-Batch Stochastic Krasnosel'ski\uı-Mann Algorithm for Nonexpansive Fixed Point Problems
von: Iiduka, Hideaki
Veröffentlicht: (2026)
von: Iiduka, Hideaki
Veröffentlicht: (2026)
Convergence Analysis of Randomized Subspace Normalized SGD under Heavy-Tailed Noise
von: Omiya, Gaku, et al.
Veröffentlicht: (2026)
von: Omiya, Gaku, et al.
Veröffentlicht: (2026)
Faster Convergence of Local SGD for Over-Parameterized Models
von: Qin, Tiancheng, et al.
Veröffentlicht: (2022)
von: Qin, Tiancheng, et al.
Veröffentlicht: (2022)
Global Convergence of SGD On Two Layer Neural Nets
von: Gopalani, Pulkit, et al.
Veröffentlicht: (2022)
von: Gopalani, Pulkit, et al.
Veröffentlicht: (2022)
High-Probability Convergence Guarantees of Decentralized SGD
von: Armacki, Aleksandar, et al.
Veröffentlicht: (2025)
von: Armacki, Aleksandar, et al.
Veröffentlicht: (2025)
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees
von: Xie, Shengping, et al.
Veröffentlicht: (2025)
von: Xie, Shengping, et al.
Veröffentlicht: (2025)
Convergence of Steepest Descent and Adam under Non-Uniform Smoothness
von: Vaswani, Sharan, et al.
Veröffentlicht: (2026)
von: Vaswani, Sharan, et al.
Veröffentlicht: (2026)
Convergence, Sticking and Escape: Stochastic Dynamics Near Critical Points in SGD
von: Dudukalov, Dmitry, et al.
Veröffentlicht: (2025)
von: Dudukalov, Dmitry, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Muon Converges under Heavy-Tailed Noise: Nonconvex Hölder-Smooth Empirical Risk Minimization
von: Iiduka, Hideaki
Veröffentlicht: (2026) -
Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
von: Umeda, Hikaru, et al.
Veröffentlicht: (2025) -
Using Stochastic Gradient Descent to Smooth Nonconvex Functions: Analysis of Implicit Graduated Optimization
von: Sato, Naoki, et al.
Veröffentlicht: (2023) -
Increasing Batch Size Improves Convergence of Stochastic Gradient Descent with Momentum
von: Kamo, Keisuke, et al.
Veröffentlicht: (2025) -
Improved Convergence Rates of Muon Optimizer for Nonconvex Optimization
von: Nagashima, Shuntaro, et al.
Veröffentlicht: (2026)