On the Interplay Between Stepsize Tuning and Progressive Sharpening
Fuente:
arXiv
Saved in:
| Main Authors: | Roulet, Vincent, Agarwala, Atish, Pedregosa, Fabian |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
High dimensional analysis reveals conservative sharpening and a stochastic edge of stability
by: Agarwala, Atish, et al.
Published: (2024)
by: Agarwala, Atish, et al.
Published: (2024)
Painless Federated Learning: An Interplay of Line-Search and Extrapolation
by: Geetika, et al.
Published: (2024)
by: Geetika, et al.
Published: (2024)
Per-example gradients: a new frontier for understanding and improving optimizers
by: Roulet, Vincent, et al.
Published: (2025)
by: Roulet, Vincent, et al.
Published: (2025)
Phases of Muon: When Muon Eclipses SignSGD
by: Paquette, Elliot, et al.
Published: (2026)
by: Paquette, Elliot, et al.
Published: (2026)
When is Momentum Extragradient Optimal? A Polynomial-Based Analysis
by: Kim, Junhyung Lyle, et al.
Published: (2022)
by: Kim, Junhyung Lyle, et al.
Published: (2022)
Bayesian Optimization for Hyperparameters Tuning in Neural Networks
by: Onorato, Gabriele
Published: (2024)
by: Onorato, Gabriele
Published: (2024)
On the Convergence of (Stochastic) Gradient Descent for Kolmogorov--Arnold Networks
by: Gao, Yihang, et al.
Published: (2024)
by: Gao, Yihang, et al.
Published: (2024)
Stepping on the Edge: Curvature Aware Learning Rate Tuners
by: Roulet, Vincent, et al.
Published: (2024)
by: Roulet, Vincent, et al.
Published: (2024)
New Hybrid Fine-Tuning Paradigm for LLMs: Algorithm Design and Convergence Analysis Framework
by: Ma, Shaocong, et al.
Published: (2026)
by: Ma, Shaocong, et al.
Published: (2026)
Score as Action: Fine-Tuning Diffusion Generative Models by Continuous-time Reinforcement Learning
by: Zhao, Hanyang, et al.
Published: (2025)
by: Zhao, Hanyang, et al.
Published: (2025)
Muon Outperforms Adam in Tail-End Associative Memory Learning
by: Wang, Shuche, et al.
Published: (2025)
by: Wang, Shuche, et al.
Published: (2025)
Stochastic Approximation with Block Coordinate Optimal Stepsizes
by: Jiang, Tao, et al.
Published: (2025)
by: Jiang, Tao, et al.
Published: (2025)
Sharpened Lazy Incremental Quasi-Newton Method
by: Lahoti, Aakash, et al.
Published: (2023)
by: Lahoti, Aakash, et al.
Published: (2023)
Constant Stepsize Q-learning: Distributional Convergence, Bias and Extrapolation
by: Zhang, Yixuan, et al.
Published: (2024)
by: Zhang, Yixuan, et al.
Published: (2024)
AdaGrad Meets Muon: Adaptive Stepsizes for Orthogonal Updates
by: Zhang, Minxin, et al.
Published: (2025)
by: Zhang, Minxin, et al.
Published: (2025)
Bias and Extrapolation in Markovian Linear Stochastic Approximation with Constant Stepsizes
by: Huo, Dongyan, et al.
Published: (2022)
by: Huo, Dongyan, et al.
Published: (2022)
New Perspectives on the Polyak Stepsize: Surrogate Functions and Negative Results
by: Orabona, Francesco, et al.
Published: (2025)
by: Orabona, Francesco, et al.
Published: (2025)
Simple Stepsize for Quasi-Newton Methods with Global Convergence Guarantees
by: Agafonov, Artem, et al.
Published: (2025)
by: Agafonov, Artem, et al.
Published: (2025)
Variance-reduced Zeroth-Order Methods for Fine-Tuning Language Models
by: Gautam, Tanmay, et al.
Published: (2024)
by: Gautam, Tanmay, et al.
Published: (2024)
On the Duality Between Sharpness-Aware Minimization and Adversarial Training
by: Zhang, Yihao, et al.
Published: (2024)
by: Zhang, Yihao, et al.
Published: (2024)
How far away are truly hyperparameter-free learning algorithms?
by: Kasimbeg, Priya, et al.
Published: (2025)
by: Kasimbeg, Priya, et al.
Published: (2025)
LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning
by: Pan, Rui, et al.
Published: (2024)
by: Pan, Rui, et al.
Published: (2024)
Understanding Forgetting in LLM Supervised Fine-Tuning and Preference Learning -- A Convex Optimization Perspective
by: Fernando, Heshan, et al.
Published: (2024)
by: Fernando, Heshan, et al.
Published: (2024)
Coupling-based Convergence Diagnostic and Stepsize Scheme for Stochastic Gradient Descent
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
An Adaptive Stochastic Gradient Method with Non-negative Gauss-Newton Stepsizes
by: Orvieto, Antonio, et al.
Published: (2024)
by: Orvieto, Antonio, et al.
Published: (2024)
Feature learning as alignment: a structural property of gradient descent in non-linear neural networks
by: Beaglehole, Daniel, et al.
Published: (2024)
by: Beaglehole, Daniel, et al.
Published: (2024)
DASA: Delay-Adaptive Multi-Agent Stochastic Approximation
by: Fabbro, Nicolò Dal, et al.
Published: (2024)
by: Fabbro, Nicolò Dal, et al.
Published: (2024)
MARINA-P: Superior Performance in Non-smooth Federated Optimization with Adaptive Stepsizes
by: Sokolov, Igor, et al.
Published: (2024)
by: Sokolov, Igor, et al.
Published: (2024)
Adaptive SGD with Line-Search and Polyak Stepsizes: Nonconvex Convergence and Accelerated Rates
by: Wu, Haotian
Published: (2025)
by: Wu, Haotian
Published: (2025)
Enhancing Stochastic Optimization for Statistical Efficiency Using ROOT-SGD with Diminishing Stepsize
by: Li, Chris Junchi
Published: (2024)
by: Li, Chris Junchi
Published: (2024)
How to escape sharp minima with random perturbations
by: Ahn, Kwangjun, et al.
Published: (2023)
by: Ahn, Kwangjun, et al.
Published: (2023)
TaskMet: Task-Driven Metric Learning for Model Learning
by: Bansal, Dishank, et al.
Published: (2023)
by: Bansal, Dishank, et al.
Published: (2023)
Multi-Objective Optimization for Sparse Deep Multi-Task Learning
by: Hotegni, S. S., et al.
Published: (2023)
by: Hotegni, S. S., et al.
Published: (2023)
Online Submodular Maximization via Online Convex Optimization
by: Salem, Tareq Si, et al.
Published: (2023)
by: Salem, Tareq Si, et al.
Published: (2023)
Comparative Analysis of Optimization Strategies for K-means Clustering in Big Data Contexts: A Review
by: Mussabayev, Ravil, et al.
Published: (2023)
by: Mussabayev, Ravil, et al.
Published: (2023)
A multiobjective continuation method to compute the regularization path of deep neural networks
by: Amakor, Augustina C., et al.
Published: (2023)
by: Amakor, Augustina C., et al.
Published: (2023)
On Finding Small Hyper-Gradients in Bilevel Optimization: Hardness Results and Improved Analysis
by: Chen, Lesi, et al.
Published: (2023)
by: Chen, Lesi, et al.
Published: (2023)
Federated Distributionally Robust Optimization with Non-Convex Objectives: Algorithm and Analysis
by: Jiao, Yang, et al.
Published: (2023)
by: Jiao, Yang, et al.
Published: (2023)
Nash Equilibria, Regularization and Computation in Optimal Transport-Based Distributionally Robust Optimization
by: Shafiee, Soroosh, et al.
Published: (2023)
by: Shafiee, Soroosh, et al.
Published: (2023)
An improved column-generation-based matheuristic for learning classification trees
by: Patel, Krunal Kishor, et al.
Published: (2023)
by: Patel, Krunal Kishor, et al.
Published: (2023)
Similar Items
-
High dimensional analysis reveals conservative sharpening and a stochastic edge of stability
by: Agarwala, Atish, et al.
Published: (2024) -
Painless Federated Learning: An Interplay of Line-Search and Extrapolation
by: Geetika, et al.
Published: (2024) -
Per-example gradients: a new frontier for understanding and improving optimizers
by: Roulet, Vincent, et al.
Published: (2025) -
Phases of Muon: When Muon Eclipses SignSGD
by: Paquette, Elliot, et al.
Published: (2026) -
When is Momentum Extragradient Optimal? A Polynomial-Based Analysis
by: Kim, Junhyung Lyle, et al.
Published: (2022)