Transformers without Normalization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Jiachen, Chen, Xinlei, He, Kaiming, LeCun, Yann, Liu, Zhuang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909648889577472
author Zhu, Jiachen
Chen, Xinlei
He, Kaiming
LeCun, Yann
Liu, Zhuang
author_facet Zhu, Jiachen
Chen, Xinlei
He, Kaiming
LeCun, Yann
Liu, Zhuang
contents Normalization layers are ubiquitous in modern neural networks and have long been considered essential. This work demonstrates that Transformers without normalization can achieve the same or better performance using a remarkably simple technique. We introduce Dynamic Tanh (DyT), an element-wise operation $DyT($x$) = \tanh(α$x$)$, as a drop-in replacement for normalization layers in Transformers. DyT is inspired by the observation that layer normalization in Transformers often produces tanh-like, $S$-shaped input-output mappings. By incorporating DyT, Transformers without normalization can match or exceed the performance of their normalized counterparts, mostly without hyperparameter tuning. We validate the effectiveness of Transformers with DyT across diverse settings, ranging from recognition to generation, supervised to self-supervised learning, and computer vision to language models. These findings challenge the conventional understanding that normalization layers are indispensable in modern neural networks, and offer new insights into their role in deep networks.
format Preprint
id arxiv_https___arxiv_org_abs_2503_10622
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Transformers without Normalization
Zhu, Jiachen
Chen, Xinlei
He, Kaiming
LeCun, Yann
Liu, Zhuang
Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Normalization layers are ubiquitous in modern neural networks and have long been considered essential. This work demonstrates that Transformers without normalization can achieve the same or better performance using a remarkably simple technique. We introduce Dynamic Tanh (DyT), an element-wise operation $DyT($x$) = \tanh(α$x$)$, as a drop-in replacement for normalization layers in Transformers. DyT is inspired by the observation that layer normalization in Transformers often produces tanh-like, $S$-shaped input-output mappings. By incorporating DyT, Transformers without normalization can match or exceed the performance of their normalized counterparts, mostly without hyperparameter tuning. We validate the effectiveness of Transformers with DyT across diverse settings, ranging from recognition to generation, supervised to self-supervised learning, and computer vision to language models. These findings challenge the conventional understanding that normalization layers are indispensable in modern neural networks, and offer new insights into their role in deep networks.
title Transformers without Normalization
topic Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.10622