What is the long-run distribution of stochastic gradient descent? A large deviations analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Azizian, Waïss, Iutzeler, Franck, Malick, Jérôme, Mertikopoulos, Panayotis |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The global convergence time of stochastic gradient descent in non-convex landscapes: Sharp estimates via large deviations
by: Azizian, Waïss, et al.
Published: (2025)
by: Azizian, Waïss, et al.
Published: (2025)
The rate of convergence of Bregman proximal methods: Local geometry vs. regularity vs. sharpness
by: Azizian, Waïss, et al.
Published: (2022)
by: Azizian, Waïss, et al.
Published: (2022)
$\texttt{skwdro}$: a library for Wasserstein distributionally robust machine learning
by: Vincent, Florian, et al.
Published: (2024)
by: Vincent, Florian, et al.
Published: (2024)
Exact Solution to Data-Driven Inverse Optimization of MILPs in Finite Time via Gradient-Based Methods
by: Kitaoka, Akira
Published: (2024)
by: Kitaoka, Akira
Published: (2024)
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization
by: Ustaomeroglu, Muhammed, et al.
Published: (2025)
by: Ustaomeroglu, Muhammed, et al.
Published: (2025)
Exponential convergence rates for momentum stochastic gradient descent in the overparametrized setting
by: Gess, Benjamin, et al.
Published: (2023)
by: Gess, Benjamin, et al.
Published: (2023)
Solving Regularized Multifacility Location Problems with Unknown Number of Centers via Difference-of-Convex Optimization
by: Geremew, W., et al.
Published: (2026)
by: Geremew, W., et al.
Published: (2026)
Kurdyka-Łojasiewicz exponent via Hadamard parametrization
by: Ouyang, Wenqing, et al.
Published: (2024)
by: Ouyang, Wenqing, et al.
Published: (2024)
Kurdyka-Łojasiewicz exponent via square transformation
by: Ouyang, Wenqing
Published: (2025)
by: Ouyang, Wenqing
Published: (2025)
Accuracy and Performance Evaluation of Quantum, Classical and Hybrid Solvers for the Max-Cut Problem
by: Vodeb, Jaka, et al.
Published: (2024)
by: Vodeb, Jaka, et al.
Published: (2024)
New results on the local-nonglobal minimizers of the generalized trust-region subproblem
by: Ai, Wenbao, et al.
Published: (2024)
by: Ai, Wenbao, et al.
Published: (2024)
Closing the duality gap of the generalized trace ratio problem
by: Yang, Meijia, et al.
Published: (2024)
by: Yang, Meijia, et al.
Published: (2024)
Optimality Conditions and Duality for Multiobjective Fractional Bilevel Optimization Problems
by: Lara, Felipe, et al.
Published: (2025)
by: Lara, Felipe, et al.
Published: (2025)
Barrier Algorithms for Constrained Non-Convex Optimization
by: Dvurechensky, Pavel, et al.
Published: (2024)
by: Dvurechensky, Pavel, et al.
Published: (2024)
Riemannian Adaptive Regularized Newton Methods with Hölder Continuous Hessians
by: Zhang, Chenyu, et al.
Published: (2023)
by: Zhang, Chenyu, et al.
Published: (2023)
An efficient proximal algorithm for squared L1 over L2 regularized sparse recovery
by: Zhang, Na, et al.
Published: (2025)
by: Zhang, Na, et al.
Published: (2025)
On Optimality Conditions for Mathematical Programming Problems Based on Strong Subdifferentials
by: Lara, Felipe, et al.
Published: (2026)
by: Lara, Felipe, et al.
Published: (2026)
Conductance Estimation in Digraphs: Submodular Transformation, Lovász Extension and Dinkelbach Iteration
by: Shao, Sihong, et al.
Published: (2025)
by: Shao, Sihong, et al.
Published: (2025)
Policy Optimization over General State and Action Spaces
by: Ju, Caleb, et al.
Published: (2022)
by: Ju, Caleb, et al.
Published: (2022)
The Geometry of Linear Program Compression: An Exact Characterization and Learning Algorithm
by: Ye, Yuhan, et al.
Published: (2026)
by: Ye, Yuhan, et al.
Published: (2026)
FSD-CAP: Fractional Subgraph Diffusion with Class-Aware Propagation for Graph Feature Imputation
by: Qiao, Xin, et al.
Published: (2026)
by: Qiao, Xin, et al.
Published: (2026)
Gaussian smoothing gradient descent for minimizing functions (GSmoothGD)
by: Starnes, Andrew, et al.
Published: (2023)
by: Starnes, Andrew, et al.
Published: (2023)
Optimization with Trained Machine Learning Models Embedded
by: Schweidtmann, Artur M., et al.
Published: (2022)
by: Schweidtmann, Artur M., et al.
Published: (2022)
Performance Estimation of second-order optimization methods on classes of univariate functions
by: Rubbens, Anne, et al.
Published: (2025)
by: Rubbens, Anne, et al.
Published: (2025)
A min-max reformulation and proximal algorithms for a class of structured nonsmooth fractional optimization problems
by: Zhou, Junpeng, et al.
Published: (2025)
by: Zhou, Junpeng, et al.
Published: (2025)
Benign landscapes for synchronization on spheres via normalized Laplacian matrices
by: McRae, Andrew D.
Published: (2025)
by: McRae, Andrew D.
Published: (2025)
Efficient Learning for Entropy-Regularized Markov Decision Processes via Multilevel Monte Carlo
by: Meunier, Matthieu, et al.
Published: (2025)
by: Meunier, Matthieu, et al.
Published: (2025)
Accelerating preconditioned ADMM via degenerate proximal point mappings
by: Sun, Defeng, et al.
Published: (2024)
by: Sun, Defeng, et al.
Published: (2024)
Uncomputability of Global Optima for Nonconvex Functions in the Oracle Model
by: Lakshmanan, K
Published: (2023)
by: Lakshmanan, K
Published: (2023)
SUDA-Muon: Structural Design Principles and Boundaries for Fully Decentralized Muon
by: Zhang, Hengrui, et al.
Published: (2026)
by: Zhang, Hengrui, et al.
Published: (2026)
Kinetic description and convergence analysis of genetic algorithms for global optimization
by: Borghi, Giacomo, et al.
Published: (2023)
by: Borghi, Giacomo, et al.
Published: (2023)
A Generalized Version of Chung's Lemma and its Applications
by: Jiang, Li, et al.
Published: (2024)
by: Jiang, Li, et al.
Published: (2024)
Stochastic Gradient Descent Revisited
by: Louzi, Azar
Published: (2024)
by: Louzi, Azar
Published: (2024)
Decentralized projected Riemannian stochastic recursive momentum method for nonconvex optimization
by: Deng, Kangkang, et al.
Published: (2024)
by: Deng, Kangkang, et al.
Published: (2024)
Minimizing Maximum Dissatisfaction in the Allocation of Indivisible Items under a Common Preference Graph
by: Chiarelli, Nina, et al.
Published: (2023)
by: Chiarelli, Nina, et al.
Published: (2023)
Adaptive Risk Mitigation in Demand Learning
by: Pakiman, Parshan, et al.
Published: (2020)
by: Pakiman, Parshan, et al.
Published: (2020)
Global Optimization of Gaussian processes
by: Schweidtmann, Artur M., et al.
Published: (2020)
by: Schweidtmann, Artur M., et al.
Published: (2020)
A Globally Optimal Portfolio for m-Sparse Sharpe Ratio Maximization
by: Lin, Yizun, et al.
Published: (2024)
by: Lin, Yizun, et al.
Published: (2024)
Machine Learning-Based Model for Postoperative Stroke Prediction in Coronary Artery Disease
by: Pan, Haonan, et al.
Published: (2025)
by: Pan, Haonan, et al.
Published: (2025)
Grassmannian optimization is NP-hard
by: Lai, Zehua, et al.
Published: (2024)
by: Lai, Zehua, et al.
Published: (2024)
Similar Items
-
The global convergence time of stochastic gradient descent in non-convex landscapes: Sharp estimates via large deviations
by: Azizian, Waïss, et al.
Published: (2025) -
The rate of convergence of Bregman proximal methods: Local geometry vs. regularity vs. sharpness
by: Azizian, Waïss, et al.
Published: (2022) -
$\texttt{skwdro}$: a library for Wasserstein distributionally robust machine learning
by: Vincent, Florian, et al.
Published: (2024) -
Exact Solution to Data-Driven Inverse Optimization of MILPs in Finite Time via Gradient-Based Methods
by: Kitaoka, Akira
Published: (2024) -
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization
by: Ustaomeroglu, Muhammed, et al.
Published: (2025)