A second-order-like optimizer with adaptive gradient scaling for deep learning
Fuente:
arXiv
Saved in:
| Main Authors: | Bolte, Jérôme, Boustany, Ryan, Pauwels, Edouard, Purica, Andrei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When majority rules, minority loses: bias amplification of gradient descent
by: Bachoc, François, et al.
Published: (2025)
by: Bachoc, François, et al.
Published: (2025)
Convergence of optimizers implies eigenvalues filtering at equilibrium
by: Bolte, Jerome, et al.
Published: (2025)
by: Bolte, Jerome, et al.
Published: (2025)
Inexact subgradient methods for semialgebraic functions
by: Bolte, Jérôme, et al.
Published: (2024)
by: Bolte, Jérôme, et al.
Published: (2024)
Fuzzy hyperparameters update in a second order optimization
by: Bensadok, Abdelaziz, et al.
Published: (2024)
by: Bensadok, Abdelaziz, et al.
Published: (2024)
Bilevel gradient methods and the Morse parametric qualification condition
by: Bolte, Jérôme, et al.
Published: (2025)
by: Bolte, Jérôme, et al.
Published: (2025)
Frugality in second-order optimization: floating-point approximations for Newton's method
by: Carrino, Giuseppe, et al.
Published: (2025)
by: Carrino, Giuseppe, et al.
Published: (2025)
Subgradient sampling for nonsmooth nonconvex minimization
by: Bolte, Jérôme, et al.
Published: (2022)
by: Bolte, Jérôme, et al.
Published: (2022)
On the numerical reliability of nonsmooth autodiff: a MaxPool case study
by: Boustany, Ryan
Published: (2024)
by: Boustany, Ryan
Published: (2024)
Derivatives of Stochastic Gradient Descent in parametric optimization
by: Iutzeler, Franck, et al.
Published: (2024)
by: Iutzeler, Franck, et al.
Published: (2024)
Learning a local trading strategy: deep reinforcement learning for grid-scale renewable energy integration
by: Ju, Caleb, et al.
Published: (2024)
by: Ju, Caleb, et al.
Published: (2024)
Bilevel reinforcement learning via the development of hyper-gradient without lower-level convexity
by: Yang, Yan, et al.
Published: (2024)
by: Yang, Yan, et al.
Published: (2024)
A note on stationarity in constrained optimization
by: Pauwels, Edouard
Published: (2024)
by: Pauwels, Edouard
Published: (2024)
Deep learning enhanced mixed integer optimization: Learning to reduce model dimensionality
by: Triantafyllou, Niki, et al.
Published: (2024)
by: Triantafyllou, Niki, et al.
Published: (2024)
Stabilizing reinforcement learning control: A modular framework for optimizing over all stable behavior
by: Lawrence, Nathan P., et al.
Published: (2023)
by: Lawrence, Nathan P., et al.
Published: (2023)
Beyond adaptive gradient: Fast-Controlled Minibatch Algorithm for large-scale optimization
by: Coppola, Corrado, et al.
Published: (2024)
by: Coppola, Corrado, et al.
Published: (2024)
Geometric and computational hardness of bilevel programming
by: Bolte, Jérôme, et al.
Published: (2024)
by: Bolte, Jérôme, et al.
Published: (2024)
Recurrent neural networks: vanishing and exploding gradients are not the end of the story
by: Zucchet, Nicolas, et al.
Published: (2024)
by: Zucchet, Nicolas, et al.
Published: (2024)
A second-order method landing on the Stiefel manifold via Newton$\unicode{x2013}$Schulz iteration
by: Xiong, Xinhui, et al.
Published: (2026)
by: Xiong, Xinhui, et al.
Published: (2026)
Tutorial on amortized optimization
by: Amos, Brandon
Published: (2022)
by: Amos, Brandon
Published: (2022)
A multiobjective continuation method to compute the regularization path of deep neural networks
by: Amakor, Augustina C., et al.
Published: (2023)
by: Amakor, Augustina C., et al.
Published: (2023)
An adaptively inexact first-order method for bilevel optimization with application to hyperparameter learning
by: Salehi, Mohammad Sadegh, et al.
Published: (2023)
by: Salehi, Mohammad Sadegh, et al.
Published: (2023)
Convergence and sample complexity of natural policy gradient primal-dual methods for constrained MDPs
by: Ding, Dongsheng, et al.
Published: (2022)
by: Ding, Dongsheng, et al.
Published: (2022)
The adjoint state method for parametric definable optimization without smoothness or uniqueness
by: Bolte, Jérôme, et al.
Published: (2026)
by: Bolte, Jérôme, et al.
Published: (2026)
A space-decoupling framework for optimization on bounded-rank matrices with orthogonally invariant constraints
by: Yang, Yan, et al.
Published: (2025)
by: Yang, Yan, et al.
Published: (2025)
Efficient and provably convergent end-to-end training of deep neural networks with linear constraints
by: Yang, Zonglin, et al.
Published: (2026)
by: Yang, Zonglin, et al.
Published: (2026)
Automated decision-making for dynamic task assignment at scale
by: Bianco, Riccardo Lo, et al.
Published: (2025)
by: Bianco, Riccardo Lo, et al.
Published: (2025)
Stochastic Optimization of Inventory at Large-scale Supply Chains
by: Jin, Zhaoyang Larry, et al.
Published: (2025)
by: Jin, Zhaoyang Larry, et al.
Published: (2025)
AdLoCo: adaptive batching significantly improves communications efficiency and convergence for Large Language Models
by: Kutuzov, Nikolay, et al.
Published: (2025)
by: Kutuzov, Nikolay, et al.
Published: (2025)
High-order expansion of Neural Ordinary Differential Equations flows
by: Izzo, Dario, et al.
Published: (2025)
by: Izzo, Dario, et al.
Published: (2025)
Wasserstein gradient flow for optimal probability measure decomposition
by: Han, Jiangze, et al.
Published: (2024)
by: Han, Jiangze, et al.
Published: (2024)
Linear attention is (maybe) all you need (to understand transformer optimization)
by: Ahn, Kwangjun, et al.
Published: (2023)
by: Ahn, Kwangjun, et al.
Published: (2023)
Optimizing the Optimizer for Physics-Informed Neural Networks and Kolmogorov-Arnold Networks
by: Kiyani, Elham, et al.
Published: (2025)
by: Kiyani, Elham, et al.
Published: (2025)
On the stability of gradient descent with second order dynamics for time-varying cost functions
by: Gibson, Travis E., et al.
Published: (2024)
by: Gibson, Travis E., et al.
Published: (2024)
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
by: Meterez, Alexandru, et al.
Published: (2025)
by: Meterez, Alexandru, et al.
Published: (2025)
An improved column-generation-based matheuristic for learning classification trees
by: Patel, Krunal Kishor, et al.
Published: (2023)
by: Patel, Krunal Kishor, et al.
Published: (2023)
High-dimensional mixed-categorical Gaussian processes with application to multidisciplinary design optimization for a green aircraft
by: Saves, Paul, et al.
Published: (2023)
by: Saves, Paul, et al.
Published: (2023)
When do spectral gradient updates help in deep learning?
by: Davis, Damek, et al.
Published: (2025)
by: Davis, Damek, et al.
Published: (2025)
Deep learning-driven scheduling algorithm for a single machine problem minimizing the total tardiness
by: Bouška, Michal, et al.
Published: (2024)
by: Bouška, Michal, et al.
Published: (2024)
Learning rate adaptive stochastic gradient descent optimization methods: numerical simulations for deep learning methods for partial differential equations and convergence analyses
by: Dereich, Steffen, et al.
Published: (2024)
by: Dereich, Steffen, et al.
Published: (2024)
On the consistent reasoning paradox of intelligence and optimal trust in AI: The power of 'I don't know'
by: Bastounis, Alexander, et al.
Published: (2024)
by: Bastounis, Alexander, et al.
Published: (2024)
Similar Items
-
When majority rules, minority loses: bias amplification of gradient descent
by: Bachoc, François, et al.
Published: (2025) -
Convergence of optimizers implies eigenvalues filtering at equilibrium
by: Bolte, Jerome, et al.
Published: (2025) -
Inexact subgradient methods for semialgebraic functions
by: Bolte, Jérôme, et al.
Published: (2024) -
Fuzzy hyperparameters update in a second order optimization
by: Bensadok, Abdelaziz, et al.
Published: (2024) -
Bilevel gradient methods and the Morse parametric qualification condition
by: Bolte, Jérôme, et al.
Published: (2025)