The Propensity for Density in Feed-forward Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913553548574720 |
|---|---|
| author | Schoots, Nandi Jackson, Alex Kholmovaia, Ali McBurney, Peter Shanahan, Murray |
| author_facet | Schoots, Nandi Jackson, Alex Kholmovaia, Ali McBurney, Peter Shanahan, Murray |
| contents | Does the process of training a neural network to solve a task tend to use all of the available weights even when the task could be solved with fewer weights? To address this question we study the effects of pruning fully connected, convolutional and residual models while varying their widths. We find that the proportion of weights that can be pruned without degrading performance is largely invariant to model size. Increasing the width of a model has little effect on the density of the pruned model relative to the increase in absolute size of the pruned network. In particular, we find substantial prunability across a large range of model sizes, where our biggest model is 50 times as wide as our smallest model. We explore three hypotheses that could explain these findings. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_14461 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | The Propensity for Density in Feed-forward Models Schoots, Nandi Jackson, Alex Kholmovaia, Ali McBurney, Peter Shanahan, Murray Machine Learning Artificial Intelligence Does the process of training a neural network to solve a task tend to use all of the available weights even when the task could be solved with fewer weights? To address this question we study the effects of pruning fully connected, convolutional and residual models while varying their widths. We find that the proportion of weights that can be pruned without degrading performance is largely invariant to model size. Increasing the width of a model has little effect on the density of the pruned model relative to the increase in absolute size of the pruned network. In particular, we find substantial prunability across a large range of model sizes, where our biggest model is 50 times as wide as our smallest model. We explore three hypotheses that could explain these findings. |
| title | The Propensity for Density in Feed-forward Models |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2410.14461 |