Efficient Neural Network Training via Subset Pretraining

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Spörer, Jan, Bermeitinger, Bernhard, Hrycej, Tomas, Limacher, Niklas, Handschuh, Siegfried
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917844686471168
author Spörer, Jan
Bermeitinger, Bernhard
Hrycej, Tomas
Limacher, Niklas
Handschuh, Siegfried
author_facet Spörer, Jan
Bermeitinger, Bernhard
Hrycej, Tomas
Limacher, Niklas
Handschuh, Siegfried
contents In training neural networks, it is common practice to use partial gradients computed over batches, mostly very small subsets of the training set. This approach is motivated by the argument that such a partial gradient is close to the true one, with precision growing only with the square root of the batch size. A theoretical justification is with the help of stochastic approximation theory. However, the conditions for the validity of this theory are not satisfied in the usual learning rate schedules. Batch processing is also difficult to combine with efficient second-order optimization methods. This proposal is based on another hypothesis: the loss minimum of the training set can be expected to be well-approximated by the minima of its subsets. Such subset minima can be computed in a fraction of the time necessary for optimizing over the whole training set. This hypothesis has been tested with the help of the MNIST, CIFAR-10, and CIFAR-100 image classification benchmarks, optionally extended by training data augmentation. The experiments have confirmed that results equivalent to conventional training can be reached. In summary, even small subsets are representative if the overdetermination ratio for the given model parameter set sufficiently exceeds unity. The computing expense can be reduced to a tenth or less.
format Preprint
id arxiv_https___arxiv_org_abs_2410_16523
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Efficient Neural Network Training via Subset Pretraining
Spörer, Jan
Bermeitinger, Bernhard
Hrycej, Tomas
Limacher, Niklas
Handschuh, Siegfried
Machine Learning
Computer Vision and Pattern Recognition
Computation
Methodology
In training neural networks, it is common practice to use partial gradients computed over batches, mostly very small subsets of the training set. This approach is motivated by the argument that such a partial gradient is close to the true one, with precision growing only with the square root of the batch size. A theoretical justification is with the help of stochastic approximation theory. However, the conditions for the validity of this theory are not satisfied in the usual learning rate schedules. Batch processing is also difficult to combine with efficient second-order optimization methods. This proposal is based on another hypothesis: the loss minimum of the training set can be expected to be well-approximated by the minima of its subsets. Such subset minima can be computed in a fraction of the time necessary for optimizing over the whole training set. This hypothesis has been tested with the help of the MNIST, CIFAR-10, and CIFAR-100 image classification benchmarks, optionally extended by training data augmentation. The experiments have confirmed that results equivalent to conventional training can be reached. In summary, even small subsets are representative if the overdetermination ratio for the given model parameter set sufficiently exceeds unity. The computing expense can be reduced to a tenth or less.
title Efficient Neural Network Training via Subset Pretraining
topic Machine Learning
Computer Vision and Pattern Recognition
Computation
Methodology
url https://arxiv.org/abs/2410.16523