Learning in Compact Spaces with Approximately Normalized Transformer

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Franke, Jörg K. H., Spiegelhalter, Urs, Nezhurina, Marianna, Jitsev, Jenia, Hutter, Frank, Hefenbrock, Michael
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917089916223488
author Franke, Jörg K. H.
Spiegelhalter, Urs
Nezhurina, Marianna
Jitsev, Jenia
Hutter, Frank
Hefenbrock, Michael
author_facet Franke, Jörg K. H.
Spiegelhalter, Urs
Nezhurina, Marianna
Jitsev, Jenia
Hutter, Frank
Hefenbrock, Michael
contents The successful training of deep neural networks requires addressing challenges such as overfitting, numerical instabilities leading to divergence, and increasing variance in the residual stream. A common solution is to apply regularization and normalization techniques that usually require tuning additional hyperparameters. An alternative is to force all parameters and representations to lie on a hypersphere. This removes the need for regularization and increases convergence speed, but comes with additional costs. In this work, we propose a more holistic, approximate normalization via simple scalar multiplications motivated by the tight concentration of the norms of high-dimensional random vectors. Additionally, instead of applying strict normalization for the parameters, we constrain their norms. These modifications remove the need for weight decay and learning rate warm-up as well, but do not increase the total number of normalization layers. Our experiments with transformer architectures show up to 40% faster convergence compared to GPT models with QK normalization, with only 3% additional runtime cost. When deriving scaling laws, we found that our method enables training with larger batch sizes while preserving the favorable scaling characteristics of classic GPT architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22014
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning in Compact Spaces with Approximately Normalized Transformer
Franke, Jörg K. H.
Spiegelhalter, Urs
Nezhurina, Marianna
Jitsev, Jenia
Hutter, Frank
Hefenbrock, Michael
Machine Learning
The successful training of deep neural networks requires addressing challenges such as overfitting, numerical instabilities leading to divergence, and increasing variance in the residual stream. A common solution is to apply regularization and normalization techniques that usually require tuning additional hyperparameters. An alternative is to force all parameters and representations to lie on a hypersphere. This removes the need for regularization and increases convergence speed, but comes with additional costs. In this work, we propose a more holistic, approximate normalization via simple scalar multiplications motivated by the tight concentration of the norms of high-dimensional random vectors. Additionally, instead of applying strict normalization for the parameters, we constrain their norms. These modifications remove the need for weight decay and learning rate warm-up as well, but do not increase the total number of normalization layers. Our experiments with transformer architectures show up to 40% faster convergence compared to GPT models with QK normalization, with only 3% additional runtime cost. When deriving scaling laws, we found that our method enables training with larger batch sizes while preserving the favorable scaling characteristics of classic GPT architectures.
title Learning in Compact Spaces with Approximately Normalized Transformer
topic Machine Learning
url https://arxiv.org/abs/2505.22014