Growing Neural Networks: Dynamic Evolution through Gradient Descent
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915412115980288 |
|---|---|
| author | Radhakrishnan, Anil Lindner, John F. Miller, Scott T. Sinha, Sudeshna Ditto, William L. |
| author_facet | Radhakrishnan, Anil Lindner, John F. Miller, Scott T. Sinha, Sudeshna Ditto, William L. |
| contents | In contrast to conventional artificial neural networks, which are structurally static, we present two approaches for evolving small networks into larger ones during training. The first method employs an auxiliary weight that directly controls network size, while the second uses a controller-generated mask to modulate neuron participation. Both approaches optimize network size through the same gradient-descent algorithm that updates the network's weights and biases. We evaluate these growing networks on nonlinear regression and classification tasks, where they consistently outperform static networks of equivalent final size. We then explore the hyperparameter space of these networks to find associated scaling relations relative to their static counterparts. Our results suggest that starting small and growing naturally may be preferable to simply starting large, particularly as neural networks continue to grow in size and energy consumption. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_18012 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Growing Neural Networks: Dynamic Evolution through Gradient Descent Radhakrishnan, Anil Lindner, John F. Miller, Scott T. Sinha, Sudeshna Ditto, William L. Machine Learning Disordered Systems and Neural Networks In contrast to conventional artificial neural networks, which are structurally static, we present two approaches for evolving small networks into larger ones during training. The first method employs an auxiliary weight that directly controls network size, while the second uses a controller-generated mask to modulate neuron participation. Both approaches optimize network size through the same gradient-descent algorithm that updates the network's weights and biases. We evaluate these growing networks on nonlinear regression and classification tasks, where they consistently outperform static networks of equivalent final size. We then explore the hyperparameter space of these networks to find associated scaling relations relative to their static counterparts. Our results suggest that starting small and growing naturally may be preferable to simply starting large, particularly as neural networks continue to grow in size and energy consumption. |
| title | Growing Neural Networks: Dynamic Evolution through Gradient Descent |
| topic | Machine Learning Disordered Systems and Neural Networks |
| url | https://arxiv.org/abs/2501.18012 |