The Propensity for Density in Feed-forward Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Schoots, Nandi, Jackson, Alex, Kholmovaia, Ali, McBurney, Peter, Shanahan, Murray
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913553548574720
author Schoots, Nandi
Jackson, Alex
Kholmovaia, Ali
McBurney, Peter
Shanahan, Murray
author_facet Schoots, Nandi
Jackson, Alex
Kholmovaia, Ali
McBurney, Peter
Shanahan, Murray
contents Does the process of training a neural network to solve a task tend to use all of the available weights even when the task could be solved with fewer weights? To address this question we study the effects of pruning fully connected, convolutional and residual models while varying their widths. We find that the proportion of weights that can be pruned without degrading performance is largely invariant to model size. Increasing the width of a model has little effect on the density of the pruned model relative to the increase in absolute size of the pruned network. In particular, we find substantial prunability across a large range of model sizes, where our biggest model is 50 times as wide as our smallest model. We explore three hypotheses that could explain these findings.
format Preprint
id arxiv_https___arxiv_org_abs_2410_14461
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The Propensity for Density in Feed-forward Models
Schoots, Nandi
Jackson, Alex
Kholmovaia, Ali
McBurney, Peter
Shanahan, Murray
Machine Learning
Artificial Intelligence
Does the process of training a neural network to solve a task tend to use all of the available weights even when the task could be solved with fewer weights? To address this question we study the effects of pruning fully connected, convolutional and residual models while varying their widths. We find that the proportion of weights that can be pruned without degrading performance is largely invariant to model size. Increasing the width of a model has little effect on the density of the pruned model relative to the increase in absolute size of the pruned network. In particular, we find substantial prunability across a large range of model sizes, where our biggest model is 50 times as wide as our smallest model. We explore three hypotheses that could explain these findings.
title The Propensity for Density in Feed-forward Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2410.14461