Parallel Layer Normalization for Universal Approximation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917258288168960 |
|---|---|
| author | Ni, Yunhao Guo, Yuxin Liu, Yuhe Sun, Wenxin Luo, Jie Wu, Wenjun Huang, Lei |
| author_facet | Ni, Yunhao Guo, Yuxin Liu, Yuhe Sun, Wenxin Luo, Jie Wu, Wenjun Huang, Lei |
| contents | This paper studies the approximation capabilities of neural networks that combine layer normalization (LN) with linear layers. We prove that networks consisting of two linear layers with parallel layer normalizations (PLNs) inserted between them (referred to as PLN-Nets) achieve universal approximation, whereas architectures that use only standard LN exhibit strictly limited expressive power.We further analyze approximation rates of shallow and deep PLN-Nets under the $L^\infty$ norm as well as in Sobolev norms. Our analysis extends beyond LN to RMSNorm, and from standard MLPs to position-wise feed-forward networks, the core building blocks used in RNNs and Transformers.Finally, we provide empirical experiments to explore other possible potentials of PLN-Nets. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_13142 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Parallel Layer Normalization for Universal Approximation Ni, Yunhao Guo, Yuxin Liu, Yuhe Sun, Wenxin Luo, Jie Wu, Wenjun Huang, Lei Machine Learning This paper studies the approximation capabilities of neural networks that combine layer normalization (LN) with linear layers. We prove that networks consisting of two linear layers with parallel layer normalizations (PLNs) inserted between them (referred to as PLN-Nets) achieve universal approximation, whereas architectures that use only standard LN exhibit strictly limited expressive power.We further analyze approximation rates of shallow and deep PLN-Nets under the $L^\infty$ norm as well as in Sobolev norms. Our analysis extends beyond LN to RMSNorm, and from standard MLPs to position-wise feed-forward networks, the core building blocks used in RNNs and Transformers.Finally, we provide empirical experiments to explore other possible potentials of PLN-Nets. |
| title | Parallel Layer Normalization for Universal Approximation |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2505.13142 |