Training neural networks faster with minimal tuning using pre-computed lists of hyperparameters for NAdamW
Fuente:
arXiv
Saved in:
| Main Authors: | Medapati, Sourabh, Kasimbeg, Priya, Krishnan, Shankar, Agarwal, Naman, Dahl, George |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How far away are truly hyperparameter-free learning algorithms?
by: Kasimbeg, Priya, et al.
Published: (2025)
by: Kasimbeg, Priya, et al.
Published: (2025)
Adaptive Gradient Methods at the Edge of Stability
by: Cohen, Jeremy M., et al.
Published: (2022)
by: Cohen, Jeremy M., et al.
Published: (2022)
What do near-optimal learning rate schedules look like?
by: Naganuma, Hiroki, et al.
Published: (2026)
by: Naganuma, Hiroki, et al.
Published: (2026)
An algorithmic framework for the optimization of deep neural networks architectures and hyperparameters
by: Keisler, Julie, et al.
Published: (2023)
by: Keisler, Julie, et al.
Published: (2023)
Accelerating Neural Network Training: An Analysis of the AlgoPerf Competition
by: Kasimbeg, Priya, et al.
Published: (2025)
by: Kasimbeg, Priya, et al.
Published: (2025)
Neural Velocity for hyperparameter tuning
by: Dalmasso, Gianluca, et al.
Published: (2025)
by: Dalmasso, Gianluca, et al.
Published: (2025)
Benchmarking Neural Network Training Algorithms
by: Dahl, George E., et al.
Published: (2023)
by: Dahl, George E., et al.
Published: (2023)
Auto Researching, not hyperparameter tuning: Convergence Analysis of 10,000 Experiments
by: Li, Xiaoyi
Published: (2026)
by: Li, Xiaoyi
Published: (2026)
QuFeX: Quantum feature extraction module for hybrid quantum-classical deep neural networks
by: Jain, Naman, et al.
Published: (2025)
by: Jain, Naman, et al.
Published: (2025)
Towards flexible perception with visual memory
by: Geirhos, Robert, et al.
Published: (2024)
by: Geirhos, Robert, et al.
Published: (2024)
Exploring possible vector systems for faster training of neural networks with preconfigured latent spaces
by: Gabdullin, Nikita
Published: (2025)
by: Gabdullin, Nikita
Published: (2025)
Training neural networks without backpropagation using particles
by: Kumar, Deepak
Published: (2024)
by: Kumar, Deepak
Published: (2024)
Be aware of overfitting by hyperparameter optimization!
by: Tetko, Igor V., et al.
Published: (2024)
by: Tetko, Igor V., et al.
Published: (2024)
Sample complexity of data-driven tuning of model hyperparameters in neural networks with structured parameter-dependent dual function
by: Balcan, Maria-Florina, et al.
Published: (2025)
by: Balcan, Maria-Florina, et al.
Published: (2025)
PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction
by: Mishra, Naman, et al.
Published: (2026)
by: Mishra, Naman, et al.
Published: (2026)
Conditional computation in neural networks: principles and research trends
by: Scardapane, Simone, et al.
Published: (2024)
by: Scardapane, Simone, et al.
Published: (2024)
Gradient Dynamics of Attention: How Cross-Entropy Sculpts Bayesian Manifolds
by: Agarwal, Naman, et al.
Published: (2025)
by: Agarwal, Naman, et al.
Published: (2025)
The Bayesian Geometry of Transformer Attention
by: Agarwal, Naman, et al.
Published: (2025)
by: Agarwal, Naman, et al.
Published: (2025)
Geometric Scaling of Bayesian Inference in LLMs
by: Agarwal, Naman, et al.
Published: (2025)
by: Agarwal, Naman, et al.
Published: (2025)
LUMOS: Large User MOdels for User Behavior Prediction
by: Nigam, Dhruv, et al.
Published: (2025)
by: Nigam, Dhruv, et al.
Published: (2025)
Application-oriented automatic hyperparameter optimization for spiking neural network prototyping
by: Fra, Vittorio
Published: (2025)
by: Fra, Vittorio
Published: (2025)
Investigating the hyperparameter space of deep neural network models for reaction coordinates
by: Kawashima, Kyohei, et al.
Published: (2024)
by: Kawashima, Kyohei, et al.
Published: (2024)
KHNNs: hypercomplex neural networks computations via Keras using TensorFlow and PyTorch
by: Niemczynowicz, Agnieszka, et al.
Published: (2024)
by: Niemczynowicz, Agnieszka, et al.
Published: (2024)
An investigation on the use of Large Language Models for hyperparameter tuning in Evolutionary Algorithms
by: Custode, Leonardo Lucio, et al.
Published: (2024)
by: Custode, Leonardo Lucio, et al.
Published: (2024)
Higher-order-ReLU-KANs (HRKANs) for solving physics-informed neural networks (PINNs) more accurately, robustly and faster
by: So, Chi Chiu, et al.
Published: (2024)
by: So, Chi Chiu, et al.
Published: (2024)
Fuzzy hyperparameters update in a second order optimization
by: Bensadok, Abdelaziz, et al.
Published: (2024)
by: Bensadok, Abdelaziz, et al.
Published: (2024)
Amortized Planning with Large-Scale Transformers: A Case Study on Chess
by: Ruoss, Anian, et al.
Published: (2024)
by: Ruoss, Anian, et al.
Published: (2024)
On the Equivalence of Regression and Classification
by: Jayadeva, et al.
Published: (2025)
by: Jayadeva, et al.
Published: (2025)
Should I try multiple optimizers when fine-tuning pre-trained Transformers for NLP tasks? Should I tune their hyperparameters?
by: Gkouti, Nefeli, et al.
Published: (2024)
by: Gkouti, Nefeli, et al.
Published: (2024)
How Does Thinking Mode Change LLM Moral Judgments? A Controlled Instant-vs-Thinking Comparison Across Five Frontier Models
by: Madur, Sai Sourabh
Published: (2026)
by: Madur, Sai Sourabh
Published: (2026)
Physics-informed neural network solves minimal surfaces in curved spacetime
by: Hashimoto, Koji, et al.
Published: (2025)
by: Hashimoto, Koji, et al.
Published: (2025)
$μ$pscaling small models: Principled warm starts and hyperparameter transfer
by: Ma, Yuxin, et al.
Published: (2026)
by: Ma, Yuxin, et al.
Published: (2026)
A multiobjective continuation method to compute the regularization path of deep neural networks
by: Amakor, Augustina C., et al.
Published: (2023)
by: Amakor, Augustina C., et al.
Published: (2023)
Phase-Aware Deep Learning with Complex-Valued CNNs for Audio Signal Applications
by: Agrawal, Naman
Published: (2025)
by: Agrawal, Naman
Published: (2025)
Target noise: A pre-training based neural network initialization for efficient high resolution learning
by: Wang, Shaowen, et al.
Published: (2026)
by: Wang, Shaowen, et al.
Published: (2026)
A comprehensive study of on-device NLP applications -- VQA, automated Form filling, Smart Replies for Linguistic Codeswitching
by: Goyal, Naman
Published: (2024)
by: Goyal, Naman
Published: (2024)
Software development effort estimation using boosting algorithms and automatic tuning of hyperparameters with Optuna
by: Maryam Hassanali, et al.
Published: (2024)
by: Maryam Hassanali, et al.
Published: (2024)
Dynamic Sparse No Training: Training-Free Fine-tuning for Sparse LLMs
by: Zhang, Yuxin, et al.
Published: (2023)
by: Zhang, Yuxin, et al.
Published: (2023)
Learning by solving differential equations
by: Dherin, Benoit, et al.
Published: (2025)
by: Dherin, Benoit, et al.
Published: (2025)
Stochastic stem bucking using mixture density neural networks
by: Schmiedel, Simon
Published: (2024)
by: Schmiedel, Simon
Published: (2024)
Similar Items
-
How far away are truly hyperparameter-free learning algorithms?
by: Kasimbeg, Priya, et al.
Published: (2025) -
Adaptive Gradient Methods at the Edge of Stability
by: Cohen, Jeremy M., et al.
Published: (2022) -
What do near-optimal learning rate schedules look like?
by: Naganuma, Hiroki, et al.
Published: (2026) -
An algorithmic framework for the optimization of deep neural networks architectures and hyperparameters
by: Keisler, Julie, et al.
Published: (2023) -
Accelerating Neural Network Training: An Analysis of the AlgoPerf Competition
by: Kasimbeg, Priya, et al.
Published: (2025)