Growing Neural Networks: Dynamic Evolution through Gradient Descent

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Radhakrishnan, Anil, Lindner, John F., Miller, Scott T., Sinha, Sudeshna, Ditto, William L.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915412115980288
author Radhakrishnan, Anil
Lindner, John F.
Miller, Scott T.
Sinha, Sudeshna
Ditto, William L.
author_facet Radhakrishnan, Anil
Lindner, John F.
Miller, Scott T.
Sinha, Sudeshna
Ditto, William L.
contents In contrast to conventional artificial neural networks, which are structurally static, we present two approaches for evolving small networks into larger ones during training. The first method employs an auxiliary weight that directly controls network size, while the second uses a controller-generated mask to modulate neuron participation. Both approaches optimize network size through the same gradient-descent algorithm that updates the network's weights and biases. We evaluate these growing networks on nonlinear regression and classification tasks, where they consistently outperform static networks of equivalent final size. We then explore the hyperparameter space of these networks to find associated scaling relations relative to their static counterparts. Our results suggest that starting small and growing naturally may be preferable to simply starting large, particularly as neural networks continue to grow in size and energy consumption.
format Preprint
id arxiv_https___arxiv_org_abs_2501_18012
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Growing Neural Networks: Dynamic Evolution through Gradient Descent
Radhakrishnan, Anil
Lindner, John F.
Miller, Scott T.
Sinha, Sudeshna
Ditto, William L.
Machine Learning
Disordered Systems and Neural Networks
In contrast to conventional artificial neural networks, which are structurally static, we present two approaches for evolving small networks into larger ones during training. The first method employs an auxiliary weight that directly controls network size, while the second uses a controller-generated mask to modulate neuron participation. Both approaches optimize network size through the same gradient-descent algorithm that updates the network's weights and biases. We evaluate these growing networks on nonlinear regression and classification tasks, where they consistently outperform static networks of equivalent final size. We then explore the hyperparameter space of these networks to find associated scaling relations relative to their static counterparts. Our results suggest that starting small and growing naturally may be preferable to simply starting large, particularly as neural networks continue to grow in size and energy consumption.
title Growing Neural Networks: Dynamic Evolution through Gradient Descent
topic Machine Learning
Disordered Systems and Neural Networks
url https://arxiv.org/abs/2501.18012