Perturbed Iterate SGD for Lipschitz Continuous Loss Functions with Numerical Error and Adaptive Step Sizes
Fuente:
arXiv
Salvato in:
| Autore principale: | Metel, Michael R. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2022
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Last-Iterate Convergence of Randomized Kaczmarz and SGD with Greedy Step Size
di: Dereziński, Michał, et al.
Pubblicazione: (2026)
di: Dereziński, Michał, et al.
Pubblicazione: (2026)
Modified K-means Algorithm with Local Optimality Guarantees
di: Li, Mingyi, et al.
Pubblicazione: (2025)
di: Li, Mingyi, et al.
Pubblicazione: (2025)
AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size
di: Ostroukhov, Petr, et al.
Pubblicazione: (2024)
di: Ostroukhov, Petr, et al.
Pubblicazione: (2024)
Fast Last-Iterate Convergence of SGD in the Smooth Interpolation Regime
di: Attia, Amit, et al.
Pubblicazione: (2025)
di: Attia, Amit, et al.
Pubblicazione: (2025)
Adaptive Step Sizes for Preconditioned Stochastic Gradient Descent
di: Köhne, Frederik, et al.
Pubblicazione: (2023)
di: Köhne, Frederik, et al.
Pubblicazione: (2023)
Randomized Feasibility Methods for Constrained Optimization with Adaptive Step Sizes
di: Chakraborty, Abhishek, et al.
Pubblicazione: (2026)
di: Chakraborty, Abhishek, et al.
Pubblicazione: (2026)
A Precise Characterization of SGD Stability Using Loss Surface Geometry
di: Dexter, Gregory, et al.
Pubblicazione: (2024)
di: Dexter, Gregory, et al.
Pubblicazione: (2024)
Global Convergence of SGD For Logistic Loss on Two Layer Neural Nets
di: Gopalani, Pulkit, et al.
Pubblicazione: (2023)
di: Gopalani, Pulkit, et al.
Pubblicazione: (2023)
Optimal Projection-Free Adaptive SGD for Matrix Optimization
di: Kovalev, Dmitry
Pubblicazione: (2026)
di: Kovalev, Dmitry
Pubblicazione: (2026)
SGD with Adaptive Preconditioning: Unified Analysis and Momentum Acceleration
di: Kovalev, Dmitry
Pubblicazione: (2025)
di: Kovalev, Dmitry
Pubblicazione: (2025)
Byzantine-Robust Distributed SGD: A Unified Analysis and Tight Error Bounds
di: Ruan, Boyuan, et al.
Pubblicazione: (2026)
di: Ruan, Boyuan, et al.
Pubblicazione: (2026)
Muon Does Not Converge on Convex Lipschitz Functions
di: Parshakova, Tetiana, et al.
Pubblicazione: (2026)
di: Parshakova, Tetiana, et al.
Pubblicazione: (2026)
Shuffling the Data, Stretching the Step-size: Sharper Bias in constant step-size SGD
di: Emmanouilidis, Konstantinos, et al.
Pubblicazione: (2026)
di: Emmanouilidis, Konstantinos, et al.
Pubblicazione: (2026)
Convergence of SGD for Training Neural Networks with Sliced Wasserstein Losses
di: Tanguy, Eloi
Pubblicazione: (2023)
di: Tanguy, Eloi
Pubblicazione: (2023)
Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
di: Umeda, Hikaru, et al.
Pubblicazione: (2025)
di: Umeda, Hikaru, et al.
Pubblicazione: (2025)
On the Convergence of DP-SGD with Adaptive Clipping
di: Shulgin, Egor, et al.
Pubblicazione: (2024)
di: Shulgin, Egor, et al.
Pubblicazione: (2024)
Adaptive SGD with Line-Search and Polyak Stepsizes: Nonconvex Convergence and Accelerated Rates
di: Wu, Haotian
Pubblicazione: (2025)
di: Wu, Haotian
Pubblicazione: (2025)
Demystifying SGD with Doubly Stochastic Gradients
di: Kim, Kyurae, et al.
Pubblicazione: (2024)
di: Kim, Kyurae, et al.
Pubblicazione: (2024)
Stochastic Weakly Convex Optimization Beyond Lipschitz Continuity
di: Gao, Wenzhi, et al.
Pubblicazione: (2024)
di: Gao, Wenzhi, et al.
Pubblicazione: (2024)
SLowcal-SGD: Slow Query Points Improve Local-SGD for Stochastic Convex Optimization
di: Dahan, Tehila, et al.
Pubblicazione: (2023)
di: Dahan, Tehila, et al.
Pubblicazione: (2023)
Revisiting Subgradient Method: Complexity and Convergence Beyond Lipschitz Continuity
di: Li, Xiao, et al.
Pubblicazione: (2023)
di: Li, Xiao, et al.
Pubblicazione: (2023)
A Convex Loss Function for Set Prediction with Optimal Trade-offs Between Size and Conditional Coverage
di: Bach, Francis
Pubblicazione: (2025)
di: Bach, Francis
Pubblicazione: (2025)
Making SGD Parameter-Free
di: Carmon, Yair, et al.
Pubblicazione: (2022)
di: Carmon, Yair, et al.
Pubblicazione: (2022)
On the Trajectories of SGD Without Replacement
di: Beneventano, Pierfrancesco
Pubblicazione: (2023)
di: Beneventano, Pierfrancesco
Pubblicazione: (2023)
An Efficient Global Optimization Algorithm with Adaptive Estimates of the Local Lipschitz Constants
di: D'Agostino, Danny
Pubblicazione: (2022)
di: D'Agostino, Danny
Pubblicazione: (2022)
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees
di: Xie, Shengping, et al.
Pubblicazione: (2025)
di: Xie, Shengping, et al.
Pubblicazione: (2025)
Shadowheart SGD: Distributed Asynchronous SGD with Optimal Time Complexity Under Arbitrary Computation and Communication Heterogeneity
di: Tyurin, Alexander, et al.
Pubblicazione: (2024)
di: Tyurin, Alexander, et al.
Pubblicazione: (2024)
Heavy-Tail Phenomenon in Decentralized SGD
di: Gurbuzbalaban, Mert, et al.
Pubblicazione: (2022)
di: Gurbuzbalaban, Mert, et al.
Pubblicazione: (2022)
Dimension-adapted Momentum Outscales SGD
di: Ferbach, Damien, et al.
Pubblicazione: (2025)
di: Ferbach, Damien, et al.
Pubblicazione: (2025)
Don't Be So Positive: Negative Step Sizes in Second-Order Methods
di: Shea, Betty, et al.
Pubblicazione: (2024)
di: Shea, Betty, et al.
Pubblicazione: (2024)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
di: Vasudeva, Bhavya, et al.
Pubblicazione: (2025)
di: Vasudeva, Bhavya, et al.
Pubblicazione: (2025)
Sign-SGD via Parameter-Free Optimization
di: Medyakov, Daniil, et al.
Pubblicazione: (2025)
di: Medyakov, Daniil, et al.
Pubblicazione: (2025)
Can SGD Handle Heavy-Tailed Noise?
di: Fatkhullin, Ilyas, et al.
Pubblicazione: (2025)
di: Fatkhullin, Ilyas, et al.
Pubblicazione: (2025)
SGD with memory: fundamental properties and stochastic acceleration
di: Yarotsky, Dmitry, et al.
Pubblicazione: (2024)
di: Yarotsky, Dmitry, et al.
Pubblicazione: (2024)
Does SGD really happen in tiny subspaces?
di: Song, Minhak, et al.
Pubblicazione: (2024)
di: Song, Minhak, et al.
Pubblicazione: (2024)
ROOT-SGD: Sharp Nonasymptotics and Near-Optimal Asymptotics in a Single Algorithm
di: Li, Chris Junchi, et al.
Pubblicazione: (2020)
di: Li, Chris Junchi, et al.
Pubblicazione: (2020)
Gradient Descent on Logistic Regression with Non-Separable Data and Large Step Sizes
di: Meng, Si Yi, et al.
Pubblicazione: (2024)
di: Meng, Si Yi, et al.
Pubblicazione: (2024)
Faster Convergence of Local SGD for Over-Parameterized Models
di: Qin, Tiancheng, et al.
Pubblicazione: (2022)
di: Qin, Tiancheng, et al.
Pubblicazione: (2022)
Global Convergence of SGD On Two Layer Neural Nets
di: Gopalani, Pulkit, et al.
Pubblicazione: (2022)
di: Gopalani, Pulkit, et al.
Pubblicazione: (2022)
Edge of Stochastic Stability: Revisiting the Edge of Stability for SGD
di: Andreyev, Arseniy, et al.
Pubblicazione: (2024)
di: Andreyev, Arseniy, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Last-Iterate Convergence of Randomized Kaczmarz and SGD with Greedy Step Size
di: Dereziński, Michał, et al.
Pubblicazione: (2026) -
Modified K-means Algorithm with Local Optimality Guarantees
di: Li, Mingyi, et al.
Pubblicazione: (2025) -
AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size
di: Ostroukhov, Petr, et al.
Pubblicazione: (2024) -
Fast Last-Iterate Convergence of SGD in the Smooth Interpolation Regime
di: Attia, Amit, et al.
Pubblicazione: (2025) -
Adaptive Step Sizes for Preconditioned Stochastic Gradient Descent
di: Köhne, Frederik, et al.
Pubblicazione: (2023)