Why Rectified Power Unit Networks Fail and How to Improve It: An Effective Field Theory Perspective

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kim, Taeyoung, Kang, Myungjoo
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912886121562112
author Kim, Taeyoung
Kang, Myungjoo
author_facet Kim, Taeyoung
Kang, Myungjoo
contents The Rectified Power Unit (RePU) activation function, a differentiable generalization of the Rectified Linear Unit (ReLU), has shown promise in constructing neural networks due to its smoothness properties. However, deep RePU networks often suffer from critical issues such as vanishing or exploding values during training, rendering them unstable regardless of hyperparameter initialization. Leveraging the perspective of effective field theory, we identify the root causes of these failures and propose the Modified Rectified Power Unit (MRePU) activation function. MRePU addresses RePU's limitations while preserving its advantages, such as differentiability and universal approximation properties. Theoretical analysis demonstrates that MRePU satisfies criticality conditions necessary for stable training, placing it in a distinct universality class. Extensive experiments validate the effectiveness of MRePU, showing significant improvements in training stability and performance across various tasks, including polynomial regression, physics-informed neural networks (PINNs) and real-world vision tasks. Our findings highlight the potential of MRePU as a robust alternative for building deep neural networks.
format Preprint
id arxiv_https___arxiv_org_abs_2408_02697
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Why Rectified Power Unit Networks Fail and How to Improve It: An Effective Field Theory Perspective
Kim, Taeyoung
Kang, Myungjoo
Machine Learning
Artificial Intelligence
The Rectified Power Unit (RePU) activation function, a differentiable generalization of the Rectified Linear Unit (ReLU), has shown promise in constructing neural networks due to its smoothness properties. However, deep RePU networks often suffer from critical issues such as vanishing or exploding values during training, rendering them unstable regardless of hyperparameter initialization. Leveraging the perspective of effective field theory, we identify the root causes of these failures and propose the Modified Rectified Power Unit (MRePU) activation function. MRePU addresses RePU's limitations while preserving its advantages, such as differentiability and universal approximation properties. Theoretical analysis demonstrates that MRePU satisfies criticality conditions necessary for stable training, placing it in a distinct universality class. Extensive experiments validate the effectiveness of MRePU, showing significant improvements in training stability and performance across various tasks, including polynomial regression, physics-informed neural networks (PINNs) and real-world vision tasks. Our findings highlight the potential of MRePU as a robust alternative for building deep neural networks.
title Why Rectified Power Unit Networks Fail and How to Improve It: An Effective Field Theory Perspective
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2408.02697