Neural networks with trainable matrix activation functions

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Zhengqi, Cao, Shuhao, Li, Yuwen, Zikatanov, Ludmil
Natura: Preprint
Pubblicazione: 2021
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912088383815680
author Liu, Zhengqi
Cao, Shuhao
Li, Yuwen
Zikatanov, Ludmil
author_facet Liu, Zhengqi
Cao, Shuhao
Li, Yuwen
Zikatanov, Ludmil
contents The training process of neural networks usually optimize weights and bias parameters of linear transformations, while nonlinear activation functions are pre-specified and fixed. This work develops a systematic approach to constructing matrix-valued activation functions whose entries are generalized from ReLU. The activation is based on matrix-vector multiplications using only scalar multiplications and comparisons. The proposed activation functions depend on parameters that are trained along with the weights and bias vectors. Neural networks based on this approach are simple and efficient and are shown to be robust in numerical experiments.
format Preprint
id arxiv_https___arxiv_org_abs_2109_09948
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle Neural networks with trainable matrix activation functions
Liu, Zhengqi
Cao, Shuhao
Li, Yuwen
Zikatanov, Ludmil
Machine Learning
The training process of neural networks usually optimize weights and bias parameters of linear transformations, while nonlinear activation functions are pre-specified and fixed. This work develops a systematic approach to constructing matrix-valued activation functions whose entries are generalized from ReLU. The activation is based on matrix-vector multiplications using only scalar multiplications and comparisons. The proposed activation functions depend on parameters that are trained along with the weights and bias vectors. Neural networks based on this approach are simple and efficient and are shown to be robust in numerical experiments.
title Neural networks with trainable matrix activation functions
topic Machine Learning
url https://arxiv.org/abs/2109.09948