When do spectral gradient updates help in deep learning?
Fuente:
arXiv
Saved in:
| Main Authors: | Davis, Damek, Drusvyatskiy, Dmitriy |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A short proof of near-linear convergence of adaptive gradient descent under fourth-order growth and convexity
by: Davis, Damek, et al.
Published: (2026)
by: Davis, Damek, et al.
Published: (2026)
Iteratively reweighted kernel machines efficiently learn sparse functions
by: Zhu, Libin, et al.
Published: (2025)
by: Zhu, Libin, et al.
Published: (2025)
Online Covariance Estimation in Nonsmooth Stochastic Approximation
by: Jiang, Liwei, et al.
Published: (2025)
by: Jiang, Liwei, et al.
Published: (2025)
Gradient descent with adaptive stepsize converges (nearly) linearly under fourth-order growth
by: Davis, Damek, et al.
Published: (2024)
by: Davis, Damek, et al.
Published: (2024)
What is the objective of reasoning with reinforcement learning?
by: Davis, Damek, et al.
Published: (2025)
by: Davis, Damek, et al.
Published: (2025)
Stochastic optimization over proximally smooth sets
by: Davis, Damek, et al.
Published: (2020)
by: Davis, Damek, et al.
Published: (2020)
Invariant Kernels: Rank Stabilization and Generalization Across Dimensions
by: Díaz, Mateo, et al.
Published: (2025)
by: Díaz, Mateo, et al.
Published: (2025)
High-dimensional Limit of SGD for Diagonal Linear Networks
by: Malaxechebarría, Begoña García, et al.
Published: (2026)
by: Malaxechebarría, Begoña García, et al.
Published: (2026)
Stochastic Approximation with Decision-Dependent Distributions: Asymptotic Normality and Optimality
by: Cutler, Joshua, et al.
Published: (2022)
by: Cutler, Joshua, et al.
Published: (2022)
Average Gradient Outer Product in kernel regression provably recovers the central subspace for multi-index models
by: Zhu, Libin, et al.
Published: (2026)
by: Zhu, Libin, et al.
Published: (2026)
A second-order-like optimizer with adaptive gradient scaling for deep learning
by: Bolte, Jérôme, et al.
Published: (2024)
by: Bolte, Jérôme, et al.
Published: (2024)
Convergence of continuous-time stochastic gradient descent with applications to deep neural networks
by: Lugosi, Gabor, et al.
Published: (2024)
by: Lugosi, Gabor, et al.
Published: (2024)
Convergence of gradient flow for learning convolutional neural networks
by: Diederen, Jona-Maria, et al.
Published: (2026)
by: Diederen, Jona-Maria, et al.
Published: (2026)
The radius of statistical efficiency
by: Cutler, Joshua, et al.
Published: (2024)
by: Cutler, Joshua, et al.
Published: (2024)
Convergence of stochastic gradient descent under a local Lojasiewicz condition for deep neural networks
by: An, Jing, et al.
Published: (2023)
by: An, Jing, et al.
Published: (2023)
Adaptive multi-gradient methods for quasiconvex vector optimization and applications to multi-task learning
by: Minh, Nguyen Anh, et al.
Published: (2024)
by: Minh, Nguyen Anh, et al.
Published: (2024)
When majority rules, minority loses: bias amplification of gradient descent
by: Bachoc, François, et al.
Published: (2025)
by: Bachoc, François, et al.
Published: (2025)
Independent policy gradient-based reinforcement learning for economic and reliable energy management of multi-microgrid systems
by: Hu, Junkai, et al.
Published: (2025)
by: Hu, Junkai, et al.
Published: (2025)
On improving generalization in a class of learning problems with the method of small parameters for weakly-controlled optimal gradient systems
by: Befekadu, Getachew K.
Published: (2024)
by: Befekadu, Getachew K.
Published: (2024)
Learning rate adaptive stochastic gradient descent optimization methods: numerical simulations for deep learning methods for partial differential equations and convergence analyses
by: Dereich, Steffen, et al.
Published: (2024)
by: Dereich, Steffen, et al.
Published: (2024)
Almost sure convergence rates of stochastic gradient methods under gradient domination
by: Weissmann, Simon, et al.
Published: (2024)
by: Weissmann, Simon, et al.
Published: (2024)
Glocal Smoothness: Line search and adaptive step sizes can help in theory too!
by: Fox, Curtis, et al.
Published: (2025)
by: Fox, Curtis, et al.
Published: (2025)
Natural Riemannian gradient for learning functional tensor networks
by: Klug, Nikolas, et al.
Published: (2026)
by: Klug, Nikolas, et al.
Published: (2026)
Beating level-set methods for 3D seismic data interpolation: a primal-dual alternating approach
by: Kumar, Rajiv, et al.
Published: (2016)
by: Kumar, Rajiv, et al.
Published: (2016)
Full error analysis of policy gradient learning algorithms for exploratory linear quadratic mean-field control problem in continuous time with common noise
by: Frikha, Noufel, et al.
Published: (2024)
by: Frikha, Noufel, et al.
Published: (2024)
Neural incomplete factorization: learning preconditioners for the conjugate gradient method
by: Häusner, Paul, et al.
Published: (2023)
by: Häusner, Paul, et al.
Published: (2023)
A stochastic gradient method for trilevel optimization
by: Giovannelli, Tommaso, et al.
Published: (2025)
by: Giovannelli, Tommaso, et al.
Published: (2025)
Generalized EXTRA stochastic gradient Langevin dynamics
by: Gurbuzbalaban, Mert, et al.
Published: (2024)
by: Gurbuzbalaban, Mert, et al.
Published: (2024)
Dealing with unbounded gradients in stochastic saddle-point optimization
by: Neu, Gergely, et al.
Published: (2024)
by: Neu, Gergely, et al.
Published: (2024)
New logarithmic step size for stochastic gradient descent
by: Shamaee, M. Soheil, et al.
Published: (2024)
by: Shamaee, M. Soheil, et al.
Published: (2024)
Scalable spectral representations for multi-agent reinforcement learning in network MDPs
by: Ren, Zhaolin, et al.
Published: (2024)
by: Ren, Zhaolin, et al.
Published: (2024)
Accelerated Methods with Complexity Separation Under Data Similarity for Federated Learning Problems
by: Bylinkin, Dmitry, et al.
Published: (2026)
by: Bylinkin, Dmitry, et al.
Published: (2026)
Problem-dependent convergence bounds for randomized linear gradient compression
by: Flynn, Thomas, et al.
Published: (2024)
by: Flynn, Thomas, et al.
Published: (2024)
Local linear convergence of gradient methods for overparameterized Gaussian mixtures
by: Wang, Jingxing, et al.
Published: (2026)
by: Wang, Jingxing, et al.
Published: (2026)
MADA: Meta-Adaptive Optimizers through hyper-gradient Descent
by: Ozkara, Kaan, et al.
Published: (2024)
by: Ozkara, Kaan, et al.
Published: (2024)
Bilevel reinforcement learning via the development of hyper-gradient without lower-level convexity
by: Yang, Yan, et al.
Published: (2024)
by: Yang, Yan, et al.
Published: (2024)
Unregularized limit of stochastic gradient method for Wasserstein distributionally robust optimization
by: Le, Tam
Published: (2025)
by: Le, Tam
Published: (2025)
Random-reshuffled SARAH does not need a full gradient computations
by: Beznosikov, Aleksandr, et al.
Published: (2021)
by: Beznosikov, Aleksandr, et al.
Published: (2021)
On the global convergence of gradient descent for wide shallow models with bounded nonlinearities
by: Petit, Romain, et al.
Published: (2026)
by: Petit, Romain, et al.
Published: (2026)
The duality structure gradient descent algorithm: analysis and applications to neural networks
by: Flynn, Thomas
Published: (2017)
by: Flynn, Thomas
Published: (2017)
Similar Items
-
A short proof of near-linear convergence of adaptive gradient descent under fourth-order growth and convexity
by: Davis, Damek, et al.
Published: (2026) -
Iteratively reweighted kernel machines efficiently learn sparse functions
by: Zhu, Libin, et al.
Published: (2025) -
Online Covariance Estimation in Nonsmooth Stochastic Approximation
by: Jiang, Liwei, et al.
Published: (2025) -
Gradient descent with adaptive stepsize converges (nearly) linearly under fourth-order growth
by: Davis, Damek, et al.
Published: (2024) -
What is the objective of reasoning with reinforcement learning?
by: Davis, Damek, et al.
Published: (2025)