Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kalra, Dayal Singh, Barkeshli, Maissam |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Why Warmup the Learning Rate? Underlying Mechanisms and Improvements
von: Kalra, Dayal Singh, et al.
Veröffentlicht: (2024)
von: Kalra, Dayal Singh, et al.
Veröffentlicht: (2024)
Universal Sharpness Dynamics in Neural Network Training: Fixed Point Analysis, Edge of Stability, and Route to Chaos
von: Kalra, Dayal Singh, et al.
Veröffentlicht: (2023)
von: Kalra, Dayal Singh, et al.
Veröffentlicht: (2023)
When Can You Get Away with Low Memory Adam?
von: Kalra, Dayal Singh, et al.
Veröffentlicht: (2025)
von: Kalra, Dayal Singh, et al.
Veröffentlicht: (2025)
On the origin of neural scaling laws: from random graphs to natural language
von: Barkeshli, Maissam, et al.
Veröffentlicht: (2026)
von: Barkeshli, Maissam, et al.
Veröffentlicht: (2026)
(How) Can Transformers Predict Pseudo-Random Numbers?
von: Tao, Tao, et al.
Veröffentlicht: (2025)
von: Tao, Tao, et al.
Veröffentlicht: (2025)
Learning Pseudorandom Numbers with Transformers: Permuted Congruential Generators, Curricula, and Interpretability
von: Tao, Tao, et al.
Veröffentlicht: (2025)
von: Tao, Tao, et al.
Veröffentlicht: (2025)
A Scalable Measure of Loss Landscape Curvature for Analyzing the Training Dynamics of LLMs
von: Kalra, Dayal Singh, et al.
Veröffentlicht: (2026)
von: Kalra, Dayal Singh, et al.
Veröffentlicht: (2026)
Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2026)
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2026)
Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
Representation Learning on a Random Lattice
von: Brill, Aryeh
Veröffentlicht: (2025)
von: Brill, Aryeh
Veröffentlicht: (2025)
Applications of Statistical Field Theory in Deep Learning
von: Ringel, Zohar, et al.
Veröffentlicht: (2025)
von: Ringel, Zohar, et al.
Veröffentlicht: (2025)
Grokking vs. Learning: Same Features, Different Encodings
von: Manning-Coe, Dmitry, et al.
Veröffentlicht: (2025)
von: Manning-Coe, Dmitry, et al.
Veröffentlicht: (2025)
Parameter Symmetry Potentially Unifies Deep Learning Theory
von: Ziyin, Liu, et al.
Veröffentlicht: (2025)
von: Ziyin, Liu, et al.
Veröffentlicht: (2025)
Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime
von: Defilippis, Leonardo, et al.
Veröffentlicht: (2025)
von: Defilippis, Leonardo, et al.
Veröffentlicht: (2025)
A Geometric Perspective on the Difficulties of Learning GNN-based SAT Solvers
von: Skenderi, Geri
Veröffentlicht: (2025)
von: Skenderi, Geri
Veröffentlicht: (2025)
Enhancing Noise-Robust Losses for Large-Scale Noisy Data Learning
von: Staats, Max, et al.
Veröffentlicht: (2023)
von: Staats, Max, et al.
Veröffentlicht: (2023)
Learning Shrinks the Hard Tail: Training-Dependent Inference Scaling in a Solvable Linear Model
von: Levi, Noam
Veröffentlicht: (2026)
von: Levi, Noam
Veröffentlicht: (2026)
Connecting NTK and NNGP: A Unified Theoretical Framework for Wide Neural Network Learning Dynamics
von: Avidan, Yehonatan, et al.
Veröffentlicht: (2023)
von: Avidan, Yehonatan, et al.
Veröffentlicht: (2023)
Predictive Coding Networks and Inference Learning: Tutorial and Survey
von: van Zwol, Björn, et al.
Veröffentlicht: (2024)
von: van Zwol, Björn, et al.
Veröffentlicht: (2024)
More Bang for the Buck: Improving the Inference of Large Language Models at a Fixed Budget using Reset and Discard (ReD)
von: Meir, Sagi, et al.
Veröffentlicht: (2026)
von: Meir, Sagi, et al.
Veröffentlicht: (2026)
Generalization through variance: how noise shapes inductive biases in diffusion models
von: Vastola, John J.
Veröffentlicht: (2025)
von: Vastola, John J.
Veröffentlicht: (2025)
Identifying internal patterns in (1+1)-dimensional directed percolation using neural networks
von: Parkhomenko, Danil, et al.
Veröffentlicht: (2025)
von: Parkhomenko, Danil, et al.
Veröffentlicht: (2025)
A Spin Glass Characterization of Neural Networks
von: Li, Jun
Veröffentlicht: (2025)
von: Li, Jun
Veröffentlicht: (2025)
A method for quantifying the generalization capabilities of generative models for solving Ising models
von: Ma, Qunlong, et al.
Veröffentlicht: (2024)
von: Ma, Qunlong, et al.
Veröffentlicht: (2024)
KAN: Kolmogorov-Arnold Networks
von: Liu, Ziming, et al.
Veröffentlicht: (2024)
von: Liu, Ziming, et al.
Veröffentlicht: (2024)
How Do Transformers "Do" Physics? Investigating the Simple Harmonic Oscillator
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2024)
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2024)
The Persian Rug: solving toy models of superposition using large-scale symmetries
von: Cowsik, Aditya, et al.
Veröffentlicht: (2024)
von: Cowsik, Aditya, et al.
Veröffentlicht: (2024)
Towards Distributed Neural Architectures
von: Cowsik, Aditya, et al.
Veröffentlicht: (2025)
von: Cowsik, Aditya, et al.
Veröffentlicht: (2025)
Smooth Kolmogorov Arnold networks enabling structural knowledge representation
von: Samadi, Moein E., et al.
Veröffentlicht: (2024)
von: Samadi, Moein E., et al.
Veröffentlicht: (2024)
Transfer Learning in $\ell_1$ Regularized Regression: Hyperparameter Selection Strategy based on Sharp Asymptotic Analysis
von: Okajima, Koki, et al.
Veröffentlicht: (2024)
von: Okajima, Koki, et al.
Veröffentlicht: (2024)
Transfer Learning in Infinite Width Feature Learning Networks
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2025)
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2025)
Predictive Coding Graphs are a Superset of Feedforward Neural Networks
von: van Zwol, Björn
Veröffentlicht: (2026)
von: van Zwol, Björn
Veröffentlicht: (2026)
Approximation Theory for Neural Networks: Old and New
von: Mukherjee, Soumendu Sundar, et al.
Veröffentlicht: (2026)
von: Mukherjee, Soumendu Sundar, et al.
Veröffentlicht: (2026)
Preisach Attention: A Hysteretic Model of Sequential Memory
von: Frydrych, Piotr
Veröffentlicht: (2026)
von: Frydrych, Piotr
Veröffentlicht: (2026)
Nature-Inspired Local Propagation
von: Betti, Alessandro, et al.
Veröffentlicht: (2024)
von: Betti, Alessandro, et al.
Veröffentlicht: (2024)
Lattice Protein Folding with Variational Annealing
von: Khandoker, Shoummo Ahsan, et al.
Veröffentlicht: (2025)
von: Khandoker, Shoummo Ahsan, et al.
Veröffentlicht: (2025)
Graph Learning Metallic Glass Discovery from Wikipedia
von: Ouyang, K. -C., et al.
Veröffentlicht: (2025)
von: Ouyang, K. -C., et al.
Veröffentlicht: (2025)
Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model
von: Bordelon, Blake, et al.
Veröffentlicht: (2026)
von: Bordelon, Blake, et al.
Veröffentlicht: (2026)
High-Dimensional Learning Dynamics of Quantized Models with Straight-Through Estimator
von: Ichikawa, Yuma, et al.
Veröffentlicht: (2025)
von: Ichikawa, Yuma, et al.
Veröffentlicht: (2025)
Dynamical Decoupling of Generalization and Overfitting in Large Two-Layer Networks
von: Montanari, Andrea, et al.
Veröffentlicht: (2025)
von: Montanari, Andrea, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Why Warmup the Learning Rate? Underlying Mechanisms and Improvements
von: Kalra, Dayal Singh, et al.
Veröffentlicht: (2024) -
Universal Sharpness Dynamics in Neural Network Training: Fixed Point Analysis, Edge of Stability, and Route to Chaos
von: Kalra, Dayal Singh, et al.
Veröffentlicht: (2023) -
When Can You Get Away with Low Memory Adam?
von: Kalra, Dayal Singh, et al.
Veröffentlicht: (2025) -
On the origin of neural scaling laws: from random graphs to natural language
von: Barkeshli, Maissam, et al.
Veröffentlicht: (2026) -
(How) Can Transformers Predict Pseudo-Random Numbers?
von: Tao, Tao, et al.
Veröffentlicht: (2025)