The Propensity for Density in Feed-forward Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Schoots, Nandi, Jackson, Alex, Kholmovaia, Ali, McBurney, Peter, Shanahan, Murray
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913553548574720
author Schoots, Nandi
Jackson, Alex
Kholmovaia, Ali
McBurney, Peter
Shanahan, Murray
author_facet Schoots, Nandi
Jackson, Alex
Kholmovaia, Ali
McBurney, Peter
Shanahan, Murray
contents Does the process of training a neural network to solve a task tend to use all of the available weights even when the task could be solved with fewer weights? To address this question we study the effects of pruning fully connected, convolutional and residual models while varying their widths. We find that the proportion of weights that can be pruned without degrading performance is largely invariant to model size. Increasing the width of a model has little effect on the density of the pruned model relative to the increase in absolute size of the pruned network. In particular, we find substantial prunability across a large range of model sizes, where our biggest model is 50 times as wide as our smallest model. We explore three hypotheses that could explain these findings.
format Preprint
id arxiv_https___arxiv_org_abs_2410_14461
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The Propensity for Density in Feed-forward Models
Schoots, Nandi
Jackson, Alex
Kholmovaia, Ali
McBurney, Peter
Shanahan, Murray
Machine Learning
Artificial Intelligence
Does the process of training a neural network to solve a task tend to use all of the available weights even when the task could be solved with fewer weights? To address this question we study the effects of pruning fully connected, convolutional and residual models while varying their widths. We find that the proportion of weights that can be pruned without degrading performance is largely invariant to model size. Increasing the width of a model has little effect on the density of the pruned model relative to the increase in absolute size of the pruned network. In particular, we find substantial prunability across a large range of model sizes, where our biggest model is 50 times as wide as our smallest model. We explore three hypotheses that could explain these findings.
title The Propensity for Density in Feed-forward Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2410.14461