Why Warmup the Learning Rate? Underlying Mechanisms and Improvements
Fuente:
arXiv
Salvato in:
| Autori principali: | Kalra, Dayal Singh, Barkeshli, Maissam |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate
di: Kalra, Dayal Singh, et al.
Pubblicazione: (2026)
di: Kalra, Dayal Singh, et al.
Pubblicazione: (2026)
Universal Sharpness Dynamics in Neural Network Training: Fixed Point Analysis, Edge of Stability, and Route to Chaos
di: Kalra, Dayal Singh, et al.
Pubblicazione: (2023)
di: Kalra, Dayal Singh, et al.
Pubblicazione: (2023)
When Can You Get Away with Low Memory Adam?
di: Kalra, Dayal Singh, et al.
Pubblicazione: (2025)
di: Kalra, Dayal Singh, et al.
Pubblicazione: (2025)
(How) Can Transformers Predict Pseudo-Random Numbers?
di: Tao, Tao, et al.
Pubblicazione: (2025)
di: Tao, Tao, et al.
Pubblicazione: (2025)
Learning Pseudorandom Numbers with Transformers: Permuted Congruential Generators, Curricula, and Interpretability
di: Tao, Tao, et al.
Pubblicazione: (2025)
di: Tao, Tao, et al.
Pubblicazione: (2025)
On the origin of neural scaling laws: from random graphs to natural language
di: Barkeshli, Maissam, et al.
Pubblicazione: (2026)
di: Barkeshli, Maissam, et al.
Pubblicazione: (2026)
A Scalable Measure of Loss Landscape Curvature for Analyzing the Training Dynamics of LLMs
di: Kalra, Dayal Singh, et al.
Pubblicazione: (2026)
di: Kalra, Dayal Singh, et al.
Pubblicazione: (2026)
Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model
di: Bordelon, Blake, et al.
Pubblicazione: (2026)
di: Bordelon, Blake, et al.
Pubblicazione: (2026)
Why Diffusion Models Don't Memorize: The Role of Implicit Dynamical Regularization in Training
di: Bonnaire, Tony, et al.
Pubblicazione: (2025)
di: Bonnaire, Tony, et al.
Pubblicazione: (2025)
Statistical Mechanics Calculations Using Variational Autoregressive Networks and Quantum Annealing
di: Tamura, Yuta, et al.
Pubblicazione: (2024)
di: Tamura, Yuta, et al.
Pubblicazione: (2024)
A Boundary-Layer Mechanism for One-Third Scaling in Online Softmax Classification
di: Kühn, Marcel, et al.
Pubblicazione: (2026)
di: Kühn, Marcel, et al.
Pubblicazione: (2026)
Transfer Learning in Infinite Width Feature Learning Networks
di: Lauditi, Clarissa, et al.
Pubblicazione: (2025)
di: Lauditi, Clarissa, et al.
Pubblicazione: (2025)
Physical Reinforcement Learning
di: Dillavou, Sam, et al.
Pubblicazione: (2025)
di: Dillavou, Sam, et al.
Pubblicazione: (2025)
Deep Learning as Neural Low-Degree Filtering: A Spectral Theory of Hierarchical Feature Learning
di: Dandi, Yatin, et al.
Pubblicazione: (2026)
di: Dandi, Yatin, et al.
Pubblicazione: (2026)
Exact Learning Dynamics of In-Context Learning in Linear Transformers and Its Application to Non-Linear Transformers
di: Mainali, Nischal, et al.
Pubblicazione: (2025)
di: Mainali, Nischal, et al.
Pubblicazione: (2025)
Hebbian Learning from First Principles
di: Albanese, Linda, et al.
Pubblicazione: (2024)
di: Albanese, Linda, et al.
Pubblicazione: (2024)
Learning and extrapolating scale-invariant processes
di: Alvez-Canepa, Anaclara, et al.
Pubblicazione: (2026)
di: Alvez-Canepa, Anaclara, et al.
Pubblicazione: (2026)
Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer
di: Lauditi, Clarissa, et al.
Pubblicazione: (2026)
di: Lauditi, Clarissa, et al.
Pubblicazione: (2026)
The effect of priors on Learning with Restricted Boltzmann Machines
di: Manzan, Gianluca, et al.
Pubblicazione: (2024)
di: Manzan, Gianluca, et al.
Pubblicazione: (2024)
Nonlocal Monte Carlo via Reinforcement Learning
di: Dobrynin, Dmitrii, et al.
Pubblicazione: (2025)
di: Dobrynin, Dmitrii, et al.
Pubblicazione: (2025)
Learning Linear Regression with Low-Rank Tasks in-Context
di: Takanami, Kaito, et al.
Pubblicazione: (2025)
di: Takanami, Kaito, et al.
Pubblicazione: (2025)
How Feature Learning Can Improve Neural Scaling Laws
di: Bordelon, Blake, et al.
Pubblicazione: (2024)
di: Bordelon, Blake, et al.
Pubblicazione: (2024)
The RL Perceptron: Generalisation Dynamics of Policy Learning in High Dimensions
di: Patel, Nishil, et al.
Pubblicazione: (2023)
di: Patel, Nishil, et al.
Pubblicazione: (2023)
The Interplay of Data Structure and Imbalance in the Learning Dynamics of Diffusion Models
di: Nicoletti, Flavio, et al.
Pubblicazione: (2026)
di: Nicoletti, Flavio, et al.
Pubblicazione: (2026)
Training Dynamics of Nonlinear Contrastive Learning Model in the High Dimensional Limit
di: Meng, Lineghuan, et al.
Pubblicazione: (2024)
di: Meng, Lineghuan, et al.
Pubblicazione: (2024)
Learning with Restricted Boltzmann Machines: Asymptotics of AMP and GD in High Dimensions
di: Xu, Yizhou, et al.
Pubblicazione: (2025)
di: Xu, Yizhou, et al.
Pubblicazione: (2025)
Disordered Dynamics in High Dimensions: Connections to Random Matrices and Machine Learning
di: Bordelon, Blake, et al.
Pubblicazione: (2026)
di: Bordelon, Blake, et al.
Pubblicazione: (2026)
Class Imbalance in Anomaly Detection: Learning from an Exactly Solvable Model
di: Pezzicoli, F. S., et al.
Pubblicazione: (2025)
di: Pezzicoli, F. S., et al.
Pubblicazione: (2025)
Modeling Structured Data Learning with Restricted Boltzmann Machines in the Teacher-Student Setting
di: Thériault, Robin, et al.
Pubblicazione: (2024)
di: Thériault, Robin, et al.
Pubblicazione: (2024)
Learning curves theory for hierarchically compositional data with power-law distributed features
di: Cagnetta, Francesco, et al.
Pubblicazione: (2025)
di: Cagnetta, Francesco, et al.
Pubblicazione: (2025)
Computational Thresholds in Multi-Modal Learning via the Spiked Matrix-Tensor Model
di: Tabanelli, Hugo, et al.
Pubblicazione: (2025)
di: Tabanelli, Hugo, et al.
Pubblicazione: (2025)
From Kernels to Features: A Multi-Scale Adaptive Theory of Feature Learning
di: Rubin, Noa, et al.
Pubblicazione: (2025)
di: Rubin, Noa, et al.
Pubblicazione: (2025)
How Deep Networks Learn Sparse and Hierarchical Data: the Sparse Random Hierarchy Model
di: Tomasini, Umberto, et al.
Pubblicazione: (2024)
di: Tomasini, Umberto, et al.
Pubblicazione: (2024)
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
di: Cagnetta, Francesco, et al.
Pubblicazione: (2025)
di: Cagnetta, Francesco, et al.
Pubblicazione: (2025)
Bilinear Sequence Regression: A Model for Learning from Long Sequences of High-dimensional Tokens
di: Erba, Vittorio, et al.
Pubblicazione: (2024)
di: Erba, Vittorio, et al.
Pubblicazione: (2024)
Statistical Physics of Deep Neural Networks: Generalization Capability, Beyond the Infinite Width, and Feature Learning
di: Ariosto, Sebastiano
Pubblicazione: (2025)
di: Ariosto, Sebastiano
Pubblicazione: (2025)
Exploring the Energy Landscape of RBMs: Reciprocal Space Insights into Bosons, Hierarchical Learning and Symmetry Breaking
di: Toledo-Marin, J. Quetzalcóatl, et al.
Pubblicazione: (2025)
di: Toledo-Marin, J. Quetzalcóatl, et al.
Pubblicazione: (2025)
Supervised Hebbian Learning
di: Alemanno, Francesco, et al.
Pubblicazione: (2022)
di: Alemanno, Francesco, et al.
Pubblicazione: (2022)
Prototype Analysis in Hopfield Networks with Hebbian Learning
di: McAlister, Hayden, et al.
Pubblicazione: (2024)
di: McAlister, Hayden, et al.
Pubblicazione: (2024)
Why is topology hard to learn?
di: Oriekhov, D. O., et al.
Pubblicazione: (2025)
di: Oriekhov, D. O., et al.
Pubblicazione: (2025)
Documenti analoghi
-
Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate
di: Kalra, Dayal Singh, et al.
Pubblicazione: (2026) -
Universal Sharpness Dynamics in Neural Network Training: Fixed Point Analysis, Edge of Stability, and Route to Chaos
di: Kalra, Dayal Singh, et al.
Pubblicazione: (2023) -
When Can You Get Away with Low Memory Adam?
di: Kalra, Dayal Singh, et al.
Pubblicazione: (2025) -
(How) Can Transformers Predict Pseudo-Random Numbers?
di: Tao, Tao, et al.
Pubblicazione: (2025) -
Learning Pseudorandom Numbers with Transformers: Permuted Congruential Generators, Curricula, and Interpretability
di: Tao, Tao, et al.
Pubblicazione: (2025)