On Suppressing Range of Adaptive Stepsizes of Adam to Improve Generalisation Performance
Fuente:
arXiv
Salvato in:
| Autore principale: | Zhang, Guoqiang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MARINA-P: Superior Performance in Non-smooth Federated Optimization with Adaptive Stepsizes
di: Sokolov, Igor, et al.
Pubblicazione: (2024)
di: Sokolov, Igor, et al.
Pubblicazione: (2024)
AdaGrad Meets Muon: Adaptive Stepsizes for Orthogonal Updates
di: Zhang, Minxin, et al.
Pubblicazione: (2025)
di: Zhang, Minxin, et al.
Pubblicazione: (2025)
Posterior Approximation using Stochastic Gradient Ascent with Adaptive Stepsize
di: Lim, Kart-Leong, et al.
Pubblicazione: (2024)
di: Lim, Kart-Leong, et al.
Pubblicazione: (2024)
Adaptive Stepsizing for Stochastic Gradient Langevin Dynamics in Bayesian Neural Networks
di: Rajpal, Rajit, et al.
Pubblicazione: (2025)
di: Rajpal, Rajit, et al.
Pubblicazione: (2025)
Minimax Optimal Convergence of Gradient Descent in Logistic Regression via Large and Adaptive Stepsizes
di: Zhang, Ruiqi, et al.
Pubblicazione: (2025)
di: Zhang, Ruiqi, et al.
Pubblicazione: (2025)
An Adaptive Stochastic Gradient Method with Non-negative Gauss-Newton Stepsizes
di: Orvieto, Antonio, et al.
Pubblicazione: (2024)
di: Orvieto, Antonio, et al.
Pubblicazione: (2024)
AdamO: A Collapse-Suppressed Optimizer for Offline RL
di: Qiao, Nan, et al.
Pubblicazione: (2026)
di: Qiao, Nan, et al.
Pubblicazione: (2026)
CaAdam: Improving Adam optimizer using connection aware methods
di: Genet, Remi, et al.
Pubblicazione: (2024)
di: Genet, Remi, et al.
Pubblicazione: (2024)
Adaptive SGD with Line-Search and Polyak Stepsizes: Nonconvex Convergence and Accelerated Rates
di: Wu, Haotian
Pubblicazione: (2025)
di: Wu, Haotian
Pubblicazione: (2025)
Large Stepsize Gradient Descent for Logistic Loss: Non-Monotonicity of the Loss Improves Optimization Efficiency
di: Wu, Jingfeng, et al.
Pubblicazione: (2024)
di: Wu, Jingfeng, et al.
Pubblicazione: (2024)
Adaptive Preconditioners Trigger Loss Spikes in Adam
di: Bai, Zhiwei, et al.
Pubblicazione: (2025)
di: Bai, Zhiwei, et al.
Pubblicazione: (2025)
Constant Stepsize Q-learning: Distributional Convergence, Bias and Extrapolation
di: Zhang, Yixuan, et al.
Pubblicazione: (2024)
di: Zhang, Yixuan, et al.
Pubblicazione: (2024)
Anon: Extrapolating Adaptivity Beyond SGD and Adam
di: Zhang, Yiheng, et al.
Pubblicazione: (2026)
di: Zhang, Yiheng, et al.
Pubblicazione: (2026)
Achieving Near-Optimal Convergence for Distributed Minimax Optimization with Adaptive Stepsizes
di: Huang, Yan, et al.
Pubblicazione: (2024)
di: Huang, Yan, et al.
Pubblicazione: (2024)
Provable Adaptivity of Adam under Non-uniform Smoothness
di: Wang, Bohan, et al.
Pubblicazione: (2022)
di: Wang, Bohan, et al.
Pubblicazione: (2022)
Stochastic Approximation with Block Coordinate Optimal Stepsizes
di: Jiang, Tao, et al.
Pubblicazione: (2025)
di: Jiang, Tao, et al.
Pubblicazione: (2025)
Large Stepsizes Accelerate Gradient Descent for Regularized Logistic Regression
di: Wu, Jingfeng, et al.
Pubblicazione: (2025)
di: Wu, Jingfeng, et al.
Pubblicazione: (2025)
Constant Stepsize Local GD for Logistic Regression: Acceleration by Instability
di: Crawshaw, Michael, et al.
Pubblicazione: (2025)
di: Crawshaw, Michael, et al.
Pubblicazione: (2025)
From Set Convergence to Pointwise Convergence: Finite-Time Guarantees for Average-Reward Q-Learning with Adaptive Stepsizes
di: Chen, Zaiwei, et al.
Pubblicazione: (2025)
di: Chen, Zaiwei, et al.
Pubblicazione: (2025)
The Collusion of Memory and Nonlinearity in Stochastic Approximation With Constant Stepsize
di: Huo, Dongyan, et al.
Pubblicazione: (2024)
di: Huo, Dongyan, et al.
Pubblicazione: (2024)
$γ$-FedHT: Stepsize-Aware Hard-Threshold Gradient Compression in Federated Learning
di: Lu, Rongwei, et al.
Pubblicazione: (2025)
di: Lu, Rongwei, et al.
Pubblicazione: (2025)
An Improved and Generalised Analysis for Spectral Clustering
di: Tyler, George, et al.
Pubblicazione: (2025)
di: Tyler, George, et al.
Pubblicazione: (2025)
Stochastic Kernel Regularisation Improves Generalisation in Deep Kernel Machines
di: Milsom, Edward, et al.
Pubblicazione: (2024)
di: Milsom, Edward, et al.
Pubblicazione: (2024)
SOAP: Improving and Stabilizing Shampoo using Adam
di: Vyas, Nikhil, et al.
Pubblicazione: (2024)
di: Vyas, Nikhil, et al.
Pubblicazione: (2024)
AdamFLIP: Adaptive Momentum Feedback Linearization Optimization for Hard Constrained PINN Training
di: Lu, Binghang, et al.
Pubblicazione: (2026)
di: Lu, Binghang, et al.
Pubblicazione: (2026)
On the Interplay Between Stepsize Tuning and Progressive Sharpening
di: Roulet, Vincent, et al.
Pubblicazione: (2023)
di: Roulet, Vincent, et al.
Pubblicazione: (2023)
Tight Bounds for Logistic Regression with Large Stepsize Gradient Descent in Low Dimension
di: Crawshaw, Michael, et al.
Pubblicazione: (2026)
di: Crawshaw, Michael, et al.
Pubblicazione: (2026)
Stepsize anything: A unified learning rate schedule for budgeted-iteration training
di: Tang, Anda, et al.
Pubblicazione: (2025)
di: Tang, Anda, et al.
Pubblicazione: (2025)
Bias and Extrapolation in Markovian Linear Stochastic Approximation with Constant Stepsizes
di: Huo, Dongyan, et al.
Pubblicazione: (2022)
di: Huo, Dongyan, et al.
Pubblicazione: (2022)
New Perspectives on the Polyak Stepsize: Surrogate Functions and Negative Results
di: Orabona, Francesco, et al.
Pubblicazione: (2025)
di: Orabona, Francesco, et al.
Pubblicazione: (2025)
Whittle Index Learning Algorithms for Restless Bandits with Constant Stepsizes
di: Mittal, Vishesh, et al.
Pubblicazione: (2024)
di: Mittal, Vishesh, et al.
Pubblicazione: (2024)
Simple Stepsize for Quasi-Newton Methods with Global Convergence Guarantees
di: Agafonov, Artem, et al.
Pubblicazione: (2025)
di: Agafonov, Artem, et al.
Pubblicazione: (2025)
HomeAdam: Adam and AdamW Algorithms Sometimes Go Home to Obtain Better Provable Generalization
di: Huang, Feihu, et al.
Pubblicazione: (2026)
di: Huang, Feihu, et al.
Pubblicazione: (2026)
Do Generalisation Results Generalise?
di: Boglioni, Matteo, et al.
Pubblicazione: (2025)
di: Boglioni, Matteo, et al.
Pubblicazione: (2025)
Adam Improves Muon: Adaptive Moment Estimation with Orthogonalized Momentum
di: Zhang, Minxin, et al.
Pubblicazione: (2026)
di: Zhang, Minxin, et al.
Pubblicazione: (2026)
Bellman Optimal Stepsize Straightening of Flow-Matching Models
di: Nguyen, Bao, et al.
Pubblicazione: (2023)
di: Nguyen, Bao, et al.
Pubblicazione: (2023)
Coupling-based Convergence Diagnostic and Stepsize Scheme for Stochastic Gradient Descent
di: Li, Xiang, et al.
Pubblicazione: (2024)
di: Li, Xiang, et al.
Pubblicazione: (2024)
How Ensembles of Distilled Policies Improve Generalisation in Reinforcement Learning
di: Weltevrede, Max, et al.
Pubblicazione: (2025)
di: Weltevrede, Max, et al.
Pubblicazione: (2025)
The Implicit Bias of Adam on Separable Data
di: Zhang, Chenyang, et al.
Pubblicazione: (2024)
di: Zhang, Chenyang, et al.
Pubblicazione: (2024)
Adam Exploits $\ell_\infty$-geometry of Loss Landscape via Coordinate-wise Adaptivity
di: Xie, Shuo, et al.
Pubblicazione: (2024)
di: Xie, Shuo, et al.
Pubblicazione: (2024)
Documenti analoghi
-
MARINA-P: Superior Performance in Non-smooth Federated Optimization with Adaptive Stepsizes
di: Sokolov, Igor, et al.
Pubblicazione: (2024) -
AdaGrad Meets Muon: Adaptive Stepsizes for Orthogonal Updates
di: Zhang, Minxin, et al.
Pubblicazione: (2025) -
Posterior Approximation using Stochastic Gradient Ascent with Adaptive Stepsize
di: Lim, Kart-Leong, et al.
Pubblicazione: (2024) -
Adaptive Stepsizing for Stochastic Gradient Langevin Dynamics in Bayesian Neural Networks
di: Rajpal, Rajit, et al.
Pubblicazione: (2025) -
Minimax Optimal Convergence of Gradient Descent in Logistic Regression via Large and Adaptive Stepsizes
di: Zhang, Ruiqi, et al.
Pubblicazione: (2025)