A framework for measuring the training efficiency of a neural architecture

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cueto-Mendoza, Eduardo, Kelleher, John D.
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917774705557504
author Cueto-Mendoza, Eduardo
Kelleher, John D.
author_facet Cueto-Mendoza, Eduardo
Kelleher, John D.
contents Measuring Efficiency in neural network system development is an open research problem. This paper presents an experimental framework to measure the training efficiency of a neural architecture. To demonstrate our approach, we analyze the training efficiency of Convolutional Neural Networks and Bayesian equivalents on the MNIST and CIFAR-10 tasks. Our results show that training efficiency decays as training progresses and varies across different stopping criteria for a given neural model and learning task. We also find a non-linear relationship between training stopping criteria, training Efficiency, model size, and training Efficiency. Furthermore, we illustrate the potential confounding effects of overtraining on measuring the training efficiency of a neural architecture. Regarding relative training efficiency across different architectures, our results indicate that CNNs are more efficient than BCNNs on both datasets. More generally, as a learning task becomes more complex, the relative difference in training efficiency between different architectures becomes more pronounced.
format Preprint
id arxiv_https___arxiv_org_abs_2409_07925
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A framework for measuring the training efficiency of a neural architecture
Cueto-Mendoza, Eduardo
Kelleher, John D.
Machine Learning
Artificial Intelligence
Measuring Efficiency in neural network system development is an open research problem. This paper presents an experimental framework to measure the training efficiency of a neural architecture. To demonstrate our approach, we analyze the training efficiency of Convolutional Neural Networks and Bayesian equivalents on the MNIST and CIFAR-10 tasks. Our results show that training efficiency decays as training progresses and varies across different stopping criteria for a given neural model and learning task. We also find a non-linear relationship between training stopping criteria, training Efficiency, model size, and training Efficiency. Furthermore, we illustrate the potential confounding effects of overtraining on measuring the training efficiency of a neural architecture. Regarding relative training efficiency across different architectures, our results indicate that CNNs are more efficient than BCNNs on both datasets. More generally, as a learning task becomes more complex, the relative difference in training efficiency between different architectures becomes more pronounced.
title A framework for measuring the training efficiency of a neural architecture
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2409.07925