How far away are truly hyperparameter-free learning algorithms?
Fuente:
arXiv
Salvato in:
| Autori principali: | Kasimbeg, Priya, Roulet, Vincent, Agarwal, Naman, Medapati, Sourabh, Pedregosa, Fabian, Agarwala, Atish, Dahl, George E. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Training neural networks faster with minimal tuning using pre-computed lists of hyperparameters for NAdamW
di: Medapati, Sourabh, et al.
Pubblicazione: (2025)
di: Medapati, Sourabh, et al.
Pubblicazione: (2025)
On the Interplay Between Stepsize Tuning and Progressive Sharpening
di: Roulet, Vincent, et al.
Pubblicazione: (2023)
di: Roulet, Vincent, et al.
Pubblicazione: (2023)
What do near-optimal learning rate schedules look like?
di: Naganuma, Hiroki, et al.
Pubblicazione: (2026)
di: Naganuma, Hiroki, et al.
Pubblicazione: (2026)
Per-example gradients: a new frontier for understanding and improving optimizers
di: Roulet, Vincent, et al.
Pubblicazione: (2025)
di: Roulet, Vincent, et al.
Pubblicazione: (2025)
Stepping on the Edge: Curvature Aware Learning Rate Tuners
di: Roulet, Vincent, et al.
Pubblicazione: (2024)
di: Roulet, Vincent, et al.
Pubblicazione: (2024)
High dimensional theory of two-phase optimizers
di: Agarwala, Atish
Pubblicazione: (2026)
di: Agarwala, Atish
Pubblicazione: (2026)
Adaptive Gradient Methods at the Edge of Stability
di: Cohen, Jeremy M., et al.
Pubblicazione: (2022)
di: Cohen, Jeremy M., et al.
Pubblicazione: (2022)
Feature learning as alignment: a structural property of gradient descent in non-linear neural networks
di: Beaglehole, Daniel, et al.
Pubblicazione: (2024)
di: Beaglehole, Daniel, et al.
Pubblicazione: (2024)
Accelerating Neural Network Training: An Analysis of the AlgoPerf Competition
di: Kasimbeg, Priya, et al.
Pubblicazione: (2025)
di: Kasimbeg, Priya, et al.
Pubblicazione: (2025)
High dimensional analysis reveals conservative sharpening and a stochastic edge of stability
di: Agarwala, Atish, et al.
Pubblicazione: (2024)
di: Agarwala, Atish, et al.
Pubblicazione: (2024)
Beyond algorithm hyperparameters: on preprocessing hyperparameters and associated pitfalls in machine learning applications
di: Sauer, Christina, et al.
Pubblicazione: (2024)
di: Sauer, Christina, et al.
Pubblicazione: (2024)
Neglected Hessian component explains mysteries in Sharpness regularization
di: Dauphin, Yann N., et al.
Pubblicazione: (2024)
di: Dauphin, Yann N., et al.
Pubblicazione: (2024)
Benchmarking Neural Network Training Algorithms
di: Dahl, George E., et al.
Pubblicazione: (2023)
di: Dahl, George E., et al.
Pubblicazione: (2023)
To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions
di: Marshall, Noah, et al.
Pubblicazione: (2024)
di: Marshall, Noah, et al.
Pubblicazione: (2024)
Exact Risk Curves of signSGD in High-Dimensions: Quantifying Preconditioning and Noise-Compression Effects
di: Xiao, Ke Liang, et al.
Pubblicazione: (2024)
di: Xiao, Ke Liang, et al.
Pubblicazione: (2024)
Avoiding spurious sharpness minimization broadens applicability of SAM
di: Singh, Sidak Pal, et al.
Pubblicazione: (2025)
di: Singh, Sidak Pal, et al.
Pubblicazione: (2025)
Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks
di: Qiu, Shikai, et al.
Pubblicazione: (2025)
di: Qiu, Shikai, et al.
Pubblicazione: (2025)
Learning by solving differential equations
di: Dherin, Benoit, et al.
Pubblicazione: (2025)
di: Dherin, Benoit, et al.
Pubblicazione: (2025)
Gradient Dynamics of Attention: How Cross-Entropy Sculpts Bayesian Manifolds
di: Agarwal, Naman, et al.
Pubblicazione: (2025)
di: Agarwal, Naman, et al.
Pubblicazione: (2025)
A Hessian-informed hyperparameter optimization for differential learning rate
di: Xu, Shiyun, et al.
Pubblicazione: (2025)
di: Xu, Shiyun, et al.
Pubblicazione: (2025)
Phases of Muon: When Muon Eclipses SignSGD
di: Paquette, Elliot, et al.
Pubblicazione: (2026)
di: Paquette, Elliot, et al.
Pubblicazione: (2026)
Better Trees: An empirical study on hyperparameter tuning of classification decision tree induction algorithms
di: Mantovani, Rafael Gomes, et al.
Pubblicazione: (2018)
di: Mantovani, Rafael Gomes, et al.
Pubblicazione: (2018)
Towards flexible perception with visual memory
di: Geirhos, Robert, et al.
Pubblicazione: (2024)
di: Geirhos, Robert, et al.
Pubblicazione: (2024)
When is Momentum Extragradient Optimal? A Polynomial-Based Analysis
di: Kim, Junhyung Lyle, et al.
Pubblicazione: (2022)
di: Kim, Junhyung Lyle, et al.
Pubblicazione: (2022)
An algorithmic framework for the optimization of deep neural networks architectures and hyperparameters
di: Keisler, Julie, et al.
Pubblicazione: (2023)
di: Keisler, Julie, et al.
Pubblicazione: (2023)
The Elements of Differentiable Programming
di: Blondel, Mathieu, et al.
Pubblicazione: (2024)
di: Blondel, Mathieu, et al.
Pubblicazione: (2024)
Towards hyperparameter-free optimization with differential privacy
di: Bu, Zhiqi, et al.
Pubblicazione: (2025)
di: Bu, Zhiqi, et al.
Pubblicazione: (2025)
Unlearning in- vs. out-of-distribution data in LLMs under gradient-based method
di: Baluta, Teodora, et al.
Pubblicazione: (2024)
di: Baluta, Teodora, et al.
Pubblicazione: (2024)
Joint Learning of Energy-based Models and their Partition Function
di: Sander, Michael E., et al.
Pubblicazione: (2025)
di: Sander, Michael E., et al.
Pubblicazione: (2025)
Be aware of overfitting by hyperparameter optimization!
di: Tetko, Igor V., et al.
Pubblicazione: (2024)
di: Tetko, Igor V., et al.
Pubblicazione: (2024)
An adaptively inexact first-order method for bilevel optimization with application to hyperparameter learning
di: Salehi, Mohammad Sadegh, et al.
Pubblicazione: (2023)
di: Salehi, Mohammad Sadegh, et al.
Pubblicazione: (2023)
How to sketch a learning algorithm
di: Gunn, Sam
Pubblicazione: (2026)
di: Gunn, Sam
Pubblicazione: (2026)
Effect of hyperparameters on variable selection in random forests
di: Fouodo, Cesaire J. K., et al.
Pubblicazione: (2023)
di: Fouodo, Cesaire J. K., et al.
Pubblicazione: (2023)
Offline-to-online hyperparameter transfer for stochastic bandits
di: Sharma, Dravyansh, et al.
Pubblicazione: (2025)
di: Sharma, Dravyansh, et al.
Pubblicazione: (2025)
Bilevel optimization for learning hyperparameters: Application to solving PDEs and inverse problems with Gaussian processes
di: Nelsen, Nicholas H., et al.
Pubblicazione: (2025)
di: Nelsen, Nicholas H., et al.
Pubblicazione: (2025)
Spectral State Space Models
di: Agarwal, Naman, et al.
Pubblicazione: (2023)
di: Agarwal, Naman, et al.
Pubblicazione: (2023)
Stacking as Accelerated Gradient Descent
di: Agarwal, Naman, et al.
Pubblicazione: (2024)
di: Agarwal, Naman, et al.
Pubblicazione: (2024)
Selecting time-series hyperparameters with the artificial jackknife
di: Pellegrino, Filippo
Pubblicazione: (2020)
di: Pellegrino, Filippo
Pubblicazione: (2020)
Calibrating dimension reduction hyperparameters in the presence of noise
di: Lin, Justin, et al.
Pubblicazione: (2023)
di: Lin, Justin, et al.
Pubblicazione: (2023)
Stability-Aware Training of Machine Learning Force Fields with Differentiable Boltzmann Estimators
di: Raja, Sanjeev, et al.
Pubblicazione: (2024)
di: Raja, Sanjeev, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Training neural networks faster with minimal tuning using pre-computed lists of hyperparameters for NAdamW
di: Medapati, Sourabh, et al.
Pubblicazione: (2025) -
On the Interplay Between Stepsize Tuning and Progressive Sharpening
di: Roulet, Vincent, et al.
Pubblicazione: (2023) -
What do near-optimal learning rate schedules look like?
di: Naganuma, Hiroki, et al.
Pubblicazione: (2026) -
Per-example gradients: a new frontier for understanding and improving optimizers
di: Roulet, Vincent, et al.
Pubblicazione: (2025) -
Stepping on the Edge: Curvature Aware Learning Rate Tuners
di: Roulet, Vincent, et al.
Pubblicazione: (2024)