ANDHRA Bandersnatch: Training Neural Networks to Predict Parallel Realities

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autor principal: Daliparthi, Venkata Satya Sai Ajay
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909408085147648
author Daliparthi, Venkata Satya Sai Ajay
author_facet Daliparthi, Venkata Satya Sai Ajay
contents Inspired by the Many-Worlds Interpretation (MWI), this work introduces a novel neural network architecture that splits the same input signal into parallel branches at each layer, utilizing a Hyper Rectified Activation, referred to as ANDHRA. The branched layers do not merge and form separate network paths, leading to multiple network heads for output prediction. For a network with a branching factor of 2 at three levels, the total number of heads is 2^3 = 8 . The individual heads are jointly trained by combining their respective loss values. However, the proposed architecture requires additional parameters and memory during training due to the additional branches. During inference, the experimental results on CIFAR-10/100 demonstrate that there exists one individual head that outperforms the baseline accuracy, achieving statistically significant improvement with equal parameters and computational cost.
format Preprint
id arxiv_https___arxiv_org_abs_2411_19213
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ANDHRA Bandersnatch: Training Neural Networks to Predict Parallel Realities
Daliparthi, Venkata Satya Sai Ajay
Computer Vision and Pattern Recognition
Inspired by the Many-Worlds Interpretation (MWI), this work introduces a novel neural network architecture that splits the same input signal into parallel branches at each layer, utilizing a Hyper Rectified Activation, referred to as ANDHRA. The branched layers do not merge and form separate network paths, leading to multiple network heads for output prediction. For a network with a branching factor of 2 at three levels, the total number of heads is 2^3 = 8 . The individual heads are jointly trained by combining their respective loss values. However, the proposed architecture requires additional parameters and memory during training due to the additional branches. During inference, the experimental results on CIFAR-10/100 demonstrate that there exists one individual head that outperforms the baseline accuracy, achieving statistically significant improvement with equal parameters and computational cost.
title ANDHRA Bandersnatch: Training Neural Networks to Predict Parallel Realities
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.19213