Don't Be So Positive: Negative Step Sizes in Second-Order Methods
Fuente:
arXiv
Saved in:
| Main Authors: | Shea, Betty, Schmidt, Mark |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Why Line Search when you can Plane Search? SO-Friendly Neural Networks allow Per-Iteration Optimization of Learning and Momentum Rates for Every Layer
by: Shea, Betty, et al.
Published: (2024)
by: Shea, Betty, et al.
Published: (2024)
Greedy Newton: Newton's Method with Exact Line Search
by: Shea, Betty, et al.
Published: (2024)
by: Shea, Betty, et al.
Published: (2024)
Don't Explain Noise: Robust Counterfactuals for Randomized Ensembles
by: Forel, Alexandre, et al.
Published: (2022)
by: Forel, Alexandre, et al.
Published: (2022)
Randomized Feasibility Methods for Constrained Optimization with Adaptive Step Sizes
by: Chakraborty, Abhishek, et al.
Published: (2026)
by: Chakraborty, Abhishek, et al.
Published: (2026)
Don't Be Greedy, Just Relax! Pruning LLMs via Frank-Wolfe
by: Roux, Christophe, et al.
Published: (2025)
by: Roux, Christophe, et al.
Published: (2025)
Better LMO-based Momentum Methods with Second-Order Information
by: Khirirat, Sarit, et al.
Published: (2025)
by: Khirirat, Sarit, et al.
Published: (2025)
First and Second Order Approximations to Stochastic Gradient Descent Methods with Momentum Terms
by: Lu, Eric
Published: (2025)
by: Lu, Eric
Published: (2025)
AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size
by: Ostroukhov, Petr, et al.
Published: (2024)
by: Ostroukhov, Petr, et al.
Published: (2024)
Can We Remove the Square-Root in Adaptive Gradient Methods? A Second-Order Perspective
by: Lin, Wu, et al.
Published: (2024)
by: Lin, Wu, et al.
Published: (2024)
Adaptive Step Sizes for Preconditioned Stochastic Gradient Descent
by: Köhne, Frederik, et al.
Published: (2023)
by: Köhne, Frederik, et al.
Published: (2023)
Krylov Cubic Regularized Newton: A Subspace Second-Order Method with Dimension-Free Convergence Rate
by: Jiang, Ruichen, et al.
Published: (2024)
by: Jiang, Ruichen, et al.
Published: (2024)
A Split-Client Approach to Second-Order Optimization
by: Chayti, El Mahdi, et al.
Published: (2025)
by: Chayti, El Mahdi, et al.
Published: (2025)
A Second-Order Majorant Algorithm for Nonnegative Matrix Factorization
by: Pham, Mai-Quyen, et al.
Published: (2023)
by: Pham, Mai-Quyen, et al.
Published: (2023)
Gradient Descent on Logistic Regression with Non-Separable Data and Large Step Sizes
by: Meng, Si Yi, et al.
Published: (2024)
by: Meng, Si Yi, et al.
Published: (2024)
AdaFisher: Adaptive Second Order Optimization via Fisher Information
by: Gomes, Damien Martins, et al.
Published: (2024)
by: Gomes, Damien Martins, et al.
Published: (2024)
Near-Optimal Distributed Minimax Optimization under the Second-Order Similarity
by: Zhou, Qihao, et al.
Published: (2024)
by: Zhou, Qihao, et al.
Published: (2024)
Gradient Descent on Logistic Regression: Do Large Step-Sizes Work with Data on the Sphere?
by: Meng, Si Yi, et al.
Published: (2025)
by: Meng, Si Yi, et al.
Published: (2025)
RanSOM: Second-Order Momentum with Randomized Scaling for Constrained and Unconstrained Optimization
by: Chayti, El Mahdi
Published: (2026)
by: Chayti, El Mahdi
Published: (2026)
Perturbed Iterate SGD for Lipschitz Continuous Loss Functions with Numerical Error and Adaptive Step Sizes
by: Metel, Michael R.
Published: (2022)
by: Metel, Michael R.
Published: (2022)
A Theoretical and Empirical Study on the Convergence of Adam with an "Exact" Constant Step Size in Non-Convex Settings
by: Mazumder, Alokendu, et al.
Published: (2023)
by: Mazumder, Alokendu, et al.
Published: (2023)
On the Role of Batch Size in Stochastic Conditional Gradient Methods
by: Islamov, Rustem, et al.
Published: (2026)
by: Islamov, Rustem, et al.
Published: (2026)
Policy Gradient with Second Order Momentum
by: Sun, Tianyu
Published: (2025)
by: Sun, Tianyu
Published: (2025)
Adaptive and Optimal Second-order Optimistic Methods for Minimax Optimization
by: Jiang, Ruichen, et al.
Published: (2024)
by: Jiang, Ruichen, et al.
Published: (2024)
Towards Practical Second-Order Optimizers in Deep Learning: Insights from Fisher Information Analysis
by: Gomes, Damien Martins
Published: (2025)
by: Gomes, Damien Martins
Published: (2025)
Safeguarded Stochastic Polyak Step Sizes for Non-smooth Optimization: Robust Performance Without Small (Sub)Gradients
by: Oikonomou, Dimitris, et al.
Published: (2025)
by: Oikonomou, Dimitris, et al.
Published: (2025)
Convex and Non-convex Federated Learning with Stale Stochastic Gradients: Diminishing Step Size is All You Need
by: Zheng, Xinran, et al.
Published: (2026)
by: Zheng, Xinran, et al.
Published: (2026)
On the Complexity of First-Order Methods in Stochastic Bilevel Optimization
by: Kwon, Jeongyeol, et al.
Published: (2024)
by: Kwon, Jeongyeol, et al.
Published: (2024)
First-Order Methods for Linearly Constrained Bilevel Optimization
by: Kornowski, Guy, et al.
Published: (2024)
by: Kornowski, Guy, et al.
Published: (2024)
Solving Convex-Concave Problems with $\tilde{\mathcal{O}}(ε^{-4/7})$ Second-Order Oracle Complexity
by: Chen, Lesi, et al.
Published: (2025)
by: Chen, Lesi, et al.
Published: (2025)
Relationship between Batch Size and Number of Steps Needed for Nonconvex Optimization of Stochastic Gradient Descent using Armijo Line Search
by: Tsukada, Yuki, et al.
Published: (2023)
by: Tsukada, Yuki, et al.
Published: (2023)
Accelerated Fully First-Order Methods for Bilevel and Minimax Optimization
by: Li, Chris Junchi
Published: (2024)
by: Li, Chris Junchi
Published: (2024)
Zeroth-Order Methods for Stochastic Nonconvex Nonsmooth Composite Optimization
by: Chen, Ziyi, et al.
Published: (2025)
by: Chen, Ziyi, et al.
Published: (2025)
Higher-Order Newton Methods with Polynomial Work per Iteration
by: Ahmadi, Amir Ali, et al.
Published: (2023)
by: Ahmadi, Amir Ali, et al.
Published: (2023)
Batched First-Order Methods for Parallel LP Solving in MIP
by: Blin, Nicolas, et al.
Published: (2026)
by: Blin, Nicolas, et al.
Published: (2026)
AdAdaGrad: Adaptive Batch Size Schemes for Adaptive Gradient Methods
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
Communication-Efficient Adaptive Batch Size Strategies for Distributed Local Gradient Methods
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
Methods with Local Steps and Random Reshuffling for Generally Smooth Non-Convex Federated Optimization
by: Demidovich, Yury, et al.
Published: (2024)
by: Demidovich, Yury, et al.
Published: (2024)
Faster Convergence of Stochastic Accelerated Gradient Descent under Interpolation
by: Mishkin, Aaron, et al.
Published: (2024)
by: Mishkin, Aaron, et al.
Published: (2024)
Implicit Bias of Spectral Descent and Muon on Multiclass Separable Data
by: Fan, Chen, et al.
Published: (2025)
by: Fan, Chen, et al.
Published: (2025)
Convergence and Sample Complexity of First-Order Methods for Agnostic Reinforcement Learning
by: Sherman, Uri, et al.
Published: (2025)
by: Sherman, Uri, et al.
Published: (2025)
Similar Items
-
Why Line Search when you can Plane Search? SO-Friendly Neural Networks allow Per-Iteration Optimization of Learning and Momentum Rates for Every Layer
by: Shea, Betty, et al.
Published: (2024) -
Greedy Newton: Newton's Method with Exact Line Search
by: Shea, Betty, et al.
Published: (2024) -
Don't Explain Noise: Robust Counterfactuals for Randomized Ensembles
by: Forel, Alexandre, et al.
Published: (2022) -
Randomized Feasibility Methods for Constrained Optimization with Adaptive Step Sizes
by: Chakraborty, Abhishek, et al.
Published: (2026) -
Don't Be Greedy, Just Relax! Pruning LLMs via Frank-Wolfe
by: Roux, Christophe, et al.
Published: (2025)