Parallel Layer Normalization for Universal Approximation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ni, Yunhao, Guo, Yuxin, Liu, Yuhe, Sun, Wenxin, Luo, Jie, Wu, Wenjun, Huang, Lei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917258288168960
author Ni, Yunhao
Guo, Yuxin
Liu, Yuhe
Sun, Wenxin
Luo, Jie
Wu, Wenjun
Huang, Lei
author_facet Ni, Yunhao
Guo, Yuxin
Liu, Yuhe
Sun, Wenxin
Luo, Jie
Wu, Wenjun
Huang, Lei
contents This paper studies the approximation capabilities of neural networks that combine layer normalization (LN) with linear layers. We prove that networks consisting of two linear layers with parallel layer normalizations (PLNs) inserted between them (referred to as PLN-Nets) achieve universal approximation, whereas architectures that use only standard LN exhibit strictly limited expressive power.We further analyze approximation rates of shallow and deep PLN-Nets under the $L^\infty$ norm as well as in Sobolev norms. Our analysis extends beyond LN to RMSNorm, and from standard MLPs to position-wise feed-forward networks, the core building blocks used in RNNs and Transformers.Finally, we provide empirical experiments to explore other possible potentials of PLN-Nets.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13142
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Parallel Layer Normalization for Universal Approximation
Ni, Yunhao
Guo, Yuxin
Liu, Yuhe
Sun, Wenxin
Luo, Jie
Wu, Wenjun
Huang, Lei
Machine Learning
This paper studies the approximation capabilities of neural networks that combine layer normalization (LN) with linear layers. We prove that networks consisting of two linear layers with parallel layer normalizations (PLNs) inserted between them (referred to as PLN-Nets) achieve universal approximation, whereas architectures that use only standard LN exhibit strictly limited expressive power.We further analyze approximation rates of shallow and deep PLN-Nets under the $L^\infty$ norm as well as in Sobolev norms. Our analysis extends beyond LN to RMSNorm, and from standard MLPs to position-wise feed-forward networks, the core building blocks used in RNNs and Transformers.Finally, we provide empirical experiments to explore other possible potentials of PLN-Nets.
title Parallel Layer Normalization for Universal Approximation
topic Machine Learning
url https://arxiv.org/abs/2505.13142