Learning Rate Transfer in Normalized Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Shigida, Boris, Hanin, Boris, Gromov, Andrey |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Memory in Optimization Algorithms Implicitly Modifies the Loss
by: Cattaneo, Matias D., et al.
Published: (2025)
by: Cattaneo, Matias D., et al.
Published: (2025)
The Effect of Mini-Batch Noise on the Implicit Bias of Adam
by: Cattaneo, Matias D., et al.
Published: (2026)
by: Cattaneo, Matias D., et al.
Published: (2026)
Implicit Bias of the JKO Scheme
by: Halmos, Peter, et al.
Published: (2025)
by: Halmos, Peter, et al.
Published: (2025)
On the Implicit Bias of Adam
by: Cattaneo, Matias D., et al.
Published: (2023)
by: Cattaneo, Matias D., et al.
Published: (2023)
Bayesian Inference with Deep Weakly Nonlinear Networks
by: Hanin, Boris, et al.
Published: (2024)
by: Hanin, Boris, et al.
Published: (2024)
nGPT: Normalized Transformer with Representation Learning on the Hypersphere
by: Loshchilov, Ilya, et al.
Published: (2024)
by: Loshchilov, Ilya, et al.
Published: (2024)
Deep Neural Nets as Hamiltonians
by: Winer, Mike, et al.
Published: (2025)
by: Winer, Mike, et al.
Published: (2025)
Quantitative CLTs in Deep Neural Networks
by: Favaro, Stefano, et al.
Published: (2023)
by: Favaro, Stefano, et al.
Published: (2023)
When Independent Sampling Outperforms Agentic Reasoning
by: Dong, Yihe, et al.
Published: (2026)
by: Dong, Yihe, et al.
Published: (2026)
Normalized Architectures are Natively 4-Bit
by: Fishman, Maxim, et al.
Published: (2026)
by: Fishman, Maxim, et al.
Published: (2026)
Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization
by: Razin, Noam, et al.
Published: (2024)
by: Razin, Noam, et al.
Published: (2024)
SVDformer: Direction-Aware Spectral Graph Embedding Learning via SVD and Transformer
by: Fang, Jiayu, et al.
Published: (2025)
by: Fang, Jiayu, et al.
Published: (2025)
Hyperparameter Transfer for Dense Associative Memories
by: Holtzman, Roi, et al.
Published: (2026)
by: Holtzman, Roi, et al.
Published: (2026)
Don't be lazy: CompleteP enables compute-efficient deep transformers
by: Dey, Nolan, et al.
Published: (2025)
by: Dey, Nolan, et al.
Published: (2025)
Les Houches Lectures on Deep Learning at Large & Infinite Width
by: Bahri, Yasaman, et al.
Published: (2023)
by: Bahri, Yasaman, et al.
Published: (2023)
Celo2: Towards Learned Optimization Free Lunch
by: Moudgil, Abhinav, et al.
Published: (2026)
by: Moudgil, Abhinav, et al.
Published: (2026)
Modified Loss of Momentum Gradient Descent: Fine-Grained Analysis
by: Cattaneo, Matias D., et al.
Published: (2025)
by: Cattaneo, Matias D., et al.
Published: (2025)
Amortized Sampling with Transferable Normalizing Flows
by: Tan, Charlie B., et al.
Published: (2025)
by: Tan, Charlie B., et al.
Published: (2025)
MODL: Multilearner Online Deep Learning
by: Valkanas, Antonios, et al.
Published: (2024)
by: Valkanas, Antonios, et al.
Published: (2024)
On the origin of neural scaling laws: from random graphs to natural language
by: Barkeshli, Maissam, et al.
Published: (2026)
by: Barkeshli, Maissam, et al.
Published: (2026)
Towards Distributed Neural Architectures
by: Cowsik, Aditya, et al.
Published: (2025)
by: Cowsik, Aditya, et al.
Published: (2025)
BARE: Leveraging Base Language Models for Few-Shot Synthetic Data Generation
by: Zhu, Alan, et al.
Published: (2025)
by: Zhu, Alan, et al.
Published: (2025)
PSTNet: Physically-Structured Turbulence Network
by: Kriuk, Boris, et al.
Published: (2026)
by: Kriuk, Boris, et al.
Published: (2026)
Are More LLM Calls All You Need? Towards Scaling Laws of Compound Inference Systems
by: Chen, Lingjiao, et al.
Published: (2024)
by: Chen, Lingjiao, et al.
Published: (2024)
Optimizing Model Selection for Compound AI Systems
by: Chen, Lingjiao, et al.
Published: (2025)
by: Chen, Lingjiao, et al.
Published: (2025)
Hyperparameter Transfer with Mixture-of-Expert Layers
by: Jiang, Tianze, et al.
Published: (2026)
by: Jiang, Tianze, et al.
Published: (2026)
Enhancing UAV Path Planning Efficiency Through Accelerated Learning
by: Viana, Joseanne, et al.
Published: (2025)
by: Viana, Joseanne, et al.
Published: (2025)
STaMP: Sequence Transformation and Mixed Precision for Low-Precision Activation Quantization
by: Federici, Marco, et al.
Published: (2025)
by: Federici, Marco, et al.
Published: (2025)
Time Transfer: On Optimal Learning Rate and Batch Size In The Infinite Data Limit
by: Filatov, Oleg, et al.
Published: (2024)
by: Filatov, Oleg, et al.
Published: (2024)
Label-Efficient Grasp Joint Prediction with Point-JEPA
by: Guzelkabaagac, Jed, et al.
Published: (2025)
by: Guzelkabaagac, Jed, et al.
Published: (2025)
DyTTP: Trajectory Prediction with Normalization-Free Transformers
by: Zhu, JianLin, et al.
Published: (2025)
by: Zhu, JianLin, et al.
Published: (2025)
Transfer Learning of Surrogate Models: Integrating Domain Warping and Affine Transformations
by: Pan, Shuaiqun, et al.
Published: (2025)
by: Pan, Shuaiqun, et al.
Published: (2025)
MARA: Continuous SE(3)-Equivariant Attention for Molecular Force Fields
by: Leonardi, Francesco, et al.
Published: (2026)
by: Leonardi, Francesco, et al.
Published: (2026)
Bayesian Inference with Shaped Deep Non-linear MLPs
by: Hanin, Boris, et al.
Published: (2026)
by: Hanin, Boris, et al.
Published: (2026)
Adaptive $Q$-Network: On-the-fly Target Selection for Deep Reinforcement Learning
by: Vincent, Théo, et al.
Published: (2024)
by: Vincent, Théo, et al.
Published: (2024)
A Proof of Learning Rate Transfer under $μ$P
by: Hayou, Soufiane
Published: (2025)
by: Hayou, Soufiane
Published: (2025)
Iterated $Q$-Network: Beyond One-Step Bellman Updates in Deep Reinforcement Learning
by: Vincent, Théo, et al.
Published: (2024)
by: Vincent, Théo, et al.
Published: (2024)
A More Realistic Evaluation of Cross-Frequency Transfer Learning and Foundation Forecasting Models
by: Olivares, Kin G., et al.
Published: (2025)
by: Olivares, Kin G., et al.
Published: (2025)
Differentiable Inductive Logic Programming for Fraud Detection
by: Wolfson, Boris, et al.
Published: (2024)
by: Wolfson, Boris, et al.
Published: (2024)
Stability of Transformers under Layer Normalization
by: Kan, Kelvin, et al.
Published: (2025)
by: Kan, Kelvin, et al.
Published: (2025)
Similar Items
-
How Memory in Optimization Algorithms Implicitly Modifies the Loss
by: Cattaneo, Matias D., et al.
Published: (2025) -
The Effect of Mini-Batch Noise on the Implicit Bias of Adam
by: Cattaneo, Matias D., et al.
Published: (2026) -
Implicit Bias of the JKO Scheme
by: Halmos, Peter, et al.
Published: (2025) -
On the Implicit Bias of Adam
by: Cattaneo, Matias D., et al.
Published: (2023) -
Bayesian Inference with Deep Weakly Nonlinear Networks
by: Hanin, Boris, et al.
Published: (2024)