Glocal Smoothness: Line search and adaptive step sizes can help in theory too!
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fox, Curtis, Mishkin, Aaron, Vaswani, Sharan, Schmidt, Mark |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Armijo Line-search Can Make (Stochastic) Gradient Descent Provably Faster
von: Vaswani, Sharan, et al.
Veröffentlicht: (2025)
von: Vaswani, Sharan, et al.
Veröffentlicht: (2025)
Convergence of Steepest Descent and Adam under Non-Uniform Smoothness
von: Vaswani, Sharan, et al.
Veröffentlicht: (2026)
von: Vaswani, Sharan, et al.
Veröffentlicht: (2026)
Towards Noise-adaptive, Problem-adaptive (Accelerated) Stochastic Gradient Descent
von: Vaswani, Sharan, et al.
Veröffentlicht: (2021)
von: Vaswani, Sharan, et al.
Veröffentlicht: (2021)
(Accelerated) Noise-adaptive Stochastic Heavy-Ball Momentum
von: Dang, Anh, et al.
Veröffentlicht: (2024)
von: Dang, Anh, et al.
Veröffentlicht: (2024)
Faster Convergence of Stochastic Accelerated Gradient Descent under Interpolation
von: Mishkin, Aaron, et al.
Veröffentlicht: (2024)
von: Mishkin, Aaron, et al.
Veröffentlicht: (2024)
Directional Smoothness and Gradient Methods: Convergence and Adaptivity
von: Mishkin, Aaron, et al.
Veröffentlicht: (2024)
von: Mishkin, Aaron, et al.
Veröffentlicht: (2024)
From Inverse Optimization to Feasibility to ERM
von: Mishra, Saurabh, et al.
Veröffentlicht: (2024)
von: Mishra, Saurabh, et al.
Veröffentlicht: (2024)
Level Set Teleportation: An Optimization Perspective
von: Mishkin, Aaron, et al.
Veröffentlicht: (2024)
von: Mishkin, Aaron, et al.
Veröffentlicht: (2024)
Why Line Search when you can Plane Search? SO-Friendly Neural Networks allow Per-Iteration Optimization of Learning and Momentum Rates for Every Layer
von: Shea, Betty, et al.
Veröffentlicht: (2024)
von: Shea, Betty, et al.
Veröffentlicht: (2024)
AltGDmin: Alternating GD and Minimization for Partly-Decoupled (Federated) Optimization
von: Vaswani, Namrata
Veröffentlicht: (2025)
von: Vaswani, Namrata
Veröffentlicht: (2025)
Sparse Polyak: an adaptive step size rule for high-dimensional M-estimation
von: Qiao, Tianqi, et al.
Veröffentlicht: (2025)
von: Qiao, Tianqi, et al.
Veröffentlicht: (2025)
Fair Supervised Learning Through Constraints on Smooth Nonconvex Unfairness-Measure Surrogates
von: Khatti, Zahra, et al.
Veröffentlicht: (2025)
von: Khatti, Zahra, et al.
Veröffentlicht: (2025)
New logarithmic step size for stochastic gradient descent
von: Shamaee, M. Soheil, et al.
Veröffentlicht: (2024)
von: Shamaee, M. Soheil, et al.
Veröffentlicht: (2024)
A Stochastic-Gradient-based Interior-Point Algorithm for Solving Smooth Bound-Constrained Optimization Problems
von: Curtis, Frank E., et al.
Veröffentlicht: (2023)
von: Curtis, Frank E., et al.
Veröffentlicht: (2023)
Shuffling the Data, Stretching the Step-size: Sharper Bias in constant step-size SGD
von: Emmanouilidis, Konstantinos, et al.
Veröffentlicht: (2026)
von: Emmanouilidis, Konstantinos, et al.
Veröffentlicht: (2026)
Convergence and concentration properties of constant step-size SGD through Markov chains
von: Merad, Ibrahim, et al.
Veröffentlicht: (2023)
von: Merad, Ibrahim, et al.
Veröffentlicht: (2023)
On the Inherent Privacy of Zeroth Order Projected Gradient Descent
von: Gupta, Devansh, et al.
Veröffentlicht: (2025)
von: Gupta, Devansh, et al.
Veröffentlicht: (2025)
Convergence rates of stochastic gradient method with independent sequences of step-size and momentum weight
von: Hwang, Wen-Liang
Veröffentlicht: (2024)
von: Hwang, Wen-Liang
Veröffentlicht: (2024)
On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization
von: Sahu, Sharan, et al.
Veröffentlicht: (2026)
von: Sahu, Sharan, et al.
Veröffentlicht: (2026)
Don't Be So Positive: Negative Step Sizes in Second-Order Methods
von: Shea, Betty, et al.
Veröffentlicht: (2024)
von: Shea, Betty, et al.
Veröffentlicht: (2024)
A theoretical and empirical study of new adaptive algorithms with additional momentum steps and shifted updates for stochastic non-convex optimization
von: Alecsa, Cristian Daniel
Veröffentlicht: (2021)
von: Alecsa, Cristian Daniel
Veröffentlicht: (2021)
Convergence of projected stochastic natural gradient variational inference for various step size and sample or batch size schedules
von: Guilmeau, Thomas, et al.
Veröffentlicht: (2026)
von: Guilmeau, Thomas, et al.
Veröffentlicht: (2026)
Integrated trucks assignment and scheduling problem with mixed service mode docks: A Q-learning based adaptive large neighborhood search algorithm
von: Li, Yueyi, et al.
Veröffentlicht: (2024)
von: Li, Yueyi, et al.
Veröffentlicht: (2024)
Smoothing the Edges: Smooth Optimization for Sparse Regularization using Hadamard Overparametrization
von: Kolb, Chris, et al.
Veröffentlicht: (2023)
von: Kolb, Chris, et al.
Veröffentlicht: (2023)
When do spectral gradient updates help in deep learning?
von: Davis, Damek, et al.
Veröffentlicht: (2025)
von: Davis, Damek, et al.
Veröffentlicht: (2025)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2025)
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2025)
Implicit Bias of Spectral Descent and Muon on Multiclass Separable Data
von: Fan, Chen, et al.
Veröffentlicht: (2025)
von: Fan, Chen, et al.
Veröffentlicht: (2025)
Smooth Quasar-Convex Optimization with Constraints
von: Martínez-Rubio, David
Veröffentlicht: (2025)
von: Martínez-Rubio, David
Veröffentlicht: (2025)
AdaGrad under Anisotropic Smoothness
von: Liu, Yuxing, et al.
Veröffentlicht: (2024)
von: Liu, Yuxing, et al.
Veröffentlicht: (2024)
Sign Operator for Coping with Heavy-Tailed Noise in Non-Convex Optimization: High Probability Bounds Under $(L_0, L_1)$-Smoothness
von: Kornilov, Nikita, et al.
Veröffentlicht: (2025)
von: Kornilov, Nikita, et al.
Veröffentlicht: (2025)
On the Interpolation Effect of Score Smoothing in Diffusion Models
von: Chen, Zhengdao
Veröffentlicht: (2025)
von: Chen, Zhengdao
Veröffentlicht: (2025)
Gradient-Normalized Smoothness for Optimization with Approximate Hessians
von: Semenov, Andrei, et al.
Veröffentlicht: (2025)
von: Semenov, Andrei, et al.
Veröffentlicht: (2025)
MGDA Converges under Generalized Smoothness, Provably
von: Zhang, Qi, et al.
Veröffentlicht: (2024)
von: Zhang, Qi, et al.
Veröffentlicht: (2024)
A Library of Mirrors: Deep Neural Nets in Low Dimensions are Convex Lasso Models with Reflection Features
von: Zeger, Emi, et al.
Veröffentlicht: (2024)
von: Zeger, Emi, et al.
Veröffentlicht: (2024)
Decentralized Stochastic Nonconvex Optimization under the Relaxed Smoothness
von: Luo, Luo, et al.
Veröffentlicht: (2025)
von: Luo, Luo, et al.
Veröffentlicht: (2025)
Revisiting Convergence: Shuffling Complexity Beyond Lipschitz Smoothness
von: He, Qi, et al.
Veröffentlicht: (2025)
von: He, Qi, et al.
Veröffentlicht: (2025)
Model approximation in MDPs with unbounded per-step cost
von: Bozkurt, Berk, et al.
Veröffentlicht: (2024)
von: Bozkurt, Berk, et al.
Veröffentlicht: (2024)
Provable Adaptivity of Adam under Non-uniform Smoothness
von: Wang, Bohan, et al.
Veröffentlicht: (2022)
von: Wang, Bohan, et al.
Veröffentlicht: (2022)
Probabilistic Smoothing with Ratio-Monotone Transforms for Global Optimization
von: Jang, Kukyoung, et al.
Veröffentlicht: (2026)
von: Jang, Kukyoung, et al.
Veröffentlicht: (2026)
Dynamic Anisotropic Smoothing for Noisy Derivative-Free Optimization
von: Reifenstein, Sam, et al.
Veröffentlicht: (2024)
von: Reifenstein, Sam, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Armijo Line-search Can Make (Stochastic) Gradient Descent Provably Faster
von: Vaswani, Sharan, et al.
Veröffentlicht: (2025) -
Convergence of Steepest Descent and Adam under Non-Uniform Smoothness
von: Vaswani, Sharan, et al.
Veröffentlicht: (2026) -
Towards Noise-adaptive, Problem-adaptive (Accelerated) Stochastic Gradient Descent
von: Vaswani, Sharan, et al.
Veröffentlicht: (2021) -
(Accelerated) Noise-adaptive Stochastic Heavy-Ball Momentum
von: Dang, Anh, et al.
Veröffentlicht: (2024) -
Faster Convergence of Stochastic Accelerated Gradient Descent under Interpolation
von: Mishkin, Aaron, et al.
Veröffentlicht: (2024)