What do near-optimal learning rate schedules look like?
Fuente:
arXiv
Saved in:
| Main Authors: | Naganuma, Hiroki, Agarwala, Atish, Kasimbeg, Priya, Dahl, George E. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How far away are truly hyperparameter-free learning algorithms?
by: Kasimbeg, Priya, et al.
Published: (2025)
by: Kasimbeg, Priya, et al.
Published: (2025)
High dimensional theory of two-phase optimizers
by: Agarwala, Atish
Published: (2026)
by: Agarwala, Atish
Published: (2026)
Per-example gradients: a new frontier for understanding and improving optimizers
by: Roulet, Vincent, et al.
Published: (2025)
by: Roulet, Vincent, et al.
Published: (2025)
Training neural networks faster with minimal tuning using pre-computed lists of hyperparameters for NAdamW
by: Medapati, Sourabh, et al.
Published: (2025)
by: Medapati, Sourabh, et al.
Published: (2025)
Feature learning as alignment: a structural property of gradient descent in non-linear neural networks
by: Beaglehole, Daniel, et al.
Published: (2024)
by: Beaglehole, Daniel, et al.
Published: (2024)
High dimensional analysis reveals conservative sharpening and a stochastic edge of stability
by: Agarwala, Atish, et al.
Published: (2024)
by: Agarwala, Atish, et al.
Published: (2024)
Neglected Hessian component explains mysteries in Sharpness regularization
by: Dauphin, Yann N., et al.
Published: (2024)
by: Dauphin, Yann N., et al.
Published: (2024)
On the Interplay Between Stepsize Tuning and Progressive Sharpening
by: Roulet, Vincent, et al.
Published: (2023)
by: Roulet, Vincent, et al.
Published: (2023)
To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions
by: Marshall, Noah, et al.
Published: (2024)
by: Marshall, Noah, et al.
Published: (2024)
Exact Risk Curves of signSGD in High-Dimensions: Quantifying Preconditioning and Noise-Compression Effects
by: Xiao, Ke Liang, et al.
Published: (2024)
by: Xiao, Ke Liang, et al.
Published: (2024)
Avoiding spurious sharpness minimization broadens applicability of SAM
by: Singh, Sidak Pal, et al.
Published: (2025)
by: Singh, Sidak Pal, et al.
Published: (2025)
Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks
by: Qiu, Shikai, et al.
Published: (2025)
by: Qiu, Shikai, et al.
Published: (2025)
Towards Understanding Variants of Invariant Risk Minimization through the Lens of Calibration
by: Yoshida, Kotaro, et al.
Published: (2024)
by: Yoshida, Kotaro, et al.
Published: (2024)
Decision-focused learning for optimal PV-Battery scheduling
by: Depoortere, Joris, et al.
Published: (2026)
by: Depoortere, Joris, et al.
Published: (2026)
Stepping on the Edge: Curvature Aware Learning Rate Tuners
by: Roulet, Vincent, et al.
Published: (2024)
by: Roulet, Vincent, et al.
Published: (2024)
Phases of Muon: When Muon Eclipses SignSGD
by: Paquette, Elliot, et al.
Published: (2026)
by: Paquette, Elliot, et al.
Published: (2026)
Geometric Insights into Focal Loss: Reducing Curvature for Enhanced Model Calibration
by: Kimura, Masanari, et al.
Published: (2024)
by: Kimura, Masanari, et al.
Published: (2024)
HyperbolicLR: Epoch insensitive learning rate scheduler
by: Kim, Tae-Geun
Published: (2024)
by: Kim, Tae-Geun
Published: (2024)
Convergence Bound and Critical Batch Size of Muon Optimizer
by: Sato, Naoki, et al.
Published: (2025)
by: Sato, Naoki, et al.
Published: (2025)
Accelerating Neural Network Training: An Analysis of the AlgoPerf Competition
by: Kasimbeg, Priya, et al.
Published: (2025)
by: Kasimbeg, Priya, et al.
Published: (2025)
Stepsize anything: A unified learning rate schedule for budgeted-iteration training
by: Tang, Anda, et al.
Published: (2025)
by: Tang, Anda, et al.
Published: (2025)
Lazy vs hasty: linearization in deep networks impacts learning schedule based on example difficulty
by: George, Thomas, et al.
Published: (2022)
by: George, Thomas, et al.
Published: (2022)
Simple and near-optimal algorithms for hidden stratification and multi-group learning
by: Tosh, Christopher, et al.
Published: (2021)
by: Tosh, Christopher, et al.
Published: (2021)
Birds look like cars: Adversarial analysis of intrinsically interpretable deep learning
by: Baniecki, Hubert, et al.
Published: (2025)
by: Baniecki, Hubert, et al.
Published: (2025)
An Empirical Study of Pre-trained Model Selection for Out-of-Distribution Generalization and Calibration
by: Naganuma, Hiroki, et al.
Published: (2023)
by: Naganuma, Hiroki, et al.
Published: (2023)
Adaptive multi-fidelity optimization with fast learning rates
by: Fiegel, Come, et al.
Published: (2026)
by: Fiegel, Come, et al.
Published: (2026)
DisTaC: Conditioning Task Vectors via Distillation for Robust Model Merging
by: Yoshida, Kotaro, et al.
Published: (2025)
by: Yoshida, Kotaro, et al.
Published: (2025)
Bayesian dynamic scheduling of multipurpose batch processes under incomplete look-ahead information
by: Zheng, Taicheng, et al.
Published: (2025)
by: Zheng, Taicheng, et al.
Published: (2025)
A Hessian-informed hyperparameter optimization for differential learning rate
by: Xu, Shiyun, et al.
Published: (2025)
by: Xu, Shiyun, et al.
Published: (2025)
Offline reinforcement learning for job-shop scheduling problems
by: Echeverria, Imanol, et al.
Published: (2024)
by: Echeverria, Imanol, et al.
Published: (2024)
This looks like what? Challenges and Future Research Directions for Part-Prototype Models
by: Elhadri, Khawla, et al.
Published: (2025)
by: Elhadri, Khawla, et al.
Published: (2025)
Dimension-free error estimate for diffusion model and optimal scheduling
by: de Bortoli, Valentin, et al.
Published: (2025)
by: de Bortoli, Valentin, et al.
Published: (2025)
Reinforcement learning-based dynamic cleaning scheduling framework for solar energy system
by: An, Heungjo
Published: (2026)
by: An, Heungjo
Published: (2026)
Takeuchi's Information Criteria as Generalization Measures for DNNs Close to NTK Regime
by: Naganuma, Hiroki, et al.
Published: (2026)
by: Naganuma, Hiroki, et al.
Published: (2026)
Distributed optimization: designed for federated learning
by: Guo, Wenyou, et al.
Published: (2025)
by: Guo, Wenyou, et al.
Published: (2025)
Another look at statistical inference with machine learning-imputed data
by: Gronsbell, Jessica, et al.
Published: (2024)
by: Gronsbell, Jessica, et al.
Published: (2024)
On Fairness of Task Arithmetic: The Role of Task Vectors
by: Naganuma, Hiroki, et al.
Published: (2025)
by: Naganuma, Hiroki, et al.
Published: (2025)
On What We Can Learn from Low-Resolution Data
by: Frehr, Theresa Dahl, et al.
Published: (2026)
by: Frehr, Theresa Dahl, et al.
Published: (2026)
An efficient deep reinforcement learning environment for flexible job-shop scheduling
by: Wu, Xinquan, et al.
Published: (2025)
by: Wu, Xinquan, et al.
Published: (2025)
Hybrid deep convolution model for lung cancer detection with transfer learning
by: Saxena, Sugandha, et al.
Published: (2025)
by: Saxena, Sugandha, et al.
Published: (2025)
Similar Items
-
How far away are truly hyperparameter-free learning algorithms?
by: Kasimbeg, Priya, et al.
Published: (2025) -
High dimensional theory of two-phase optimizers
by: Agarwala, Atish
Published: (2026) -
Per-example gradients: a new frontier for understanding and improving optimizers
by: Roulet, Vincent, et al.
Published: (2025) -
Training neural networks faster with minimal tuning using pre-computed lists of hyperparameters for NAdamW
by: Medapati, Sourabh, et al.
Published: (2025) -
Feature learning as alignment: a structural property of gradient descent in non-linear neural networks
by: Beaglehole, Daniel, et al.
Published: (2024)