Growth strategies for arbitrary DAG neural architectures

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Douka, Stella, Verbockhaven, Manon, Rudkiewicz, Théo, Rivaud, Stéphane, Landes, François P., Chevallier, Sylvain, Charpiat, Guillaume
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912231914995712
author Douka, Stella
Verbockhaven, Manon
Rudkiewicz, Théo
Rivaud, Stéphane
Landes, François P.
Chevallier, Sylvain
Charpiat, Guillaume
author_facet Douka, Stella
Verbockhaven, Manon
Rudkiewicz, Théo
Rivaud, Stéphane
Landes, François P.
Chevallier, Sylvain
Charpiat, Guillaume
contents Deep learning has shown impressive results obtained at the cost of training huge neural networks. However, the larger the architecture, the higher the computational, financial, and environmental costs during training and inference. We aim at reducing both training and inference durations. We focus on Neural Architecture Growth, which can increase the size of a small model when needed, directly during training using information from the backpropagation. We expand existing work and freely grow neural networks in the form of any Directed Acyclic Graph by reducing expressivity bottlenecks in the architecture. We explore strategies to reduce excessive computations and steer network growth toward more parameter-efficient architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2501_12690
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Growth strategies for arbitrary DAG neural architectures
Douka, Stella
Verbockhaven, Manon
Rudkiewicz, Théo
Rivaud, Stéphane
Landes, François P.
Chevallier, Sylvain
Charpiat, Guillaume
Machine Learning
Artificial Intelligence
Deep learning has shown impressive results obtained at the cost of training huge neural networks. However, the larger the architecture, the higher the computational, financial, and environmental costs during training and inference. We aim at reducing both training and inference durations. We focus on Neural Architecture Growth, which can increase the size of a small model when needed, directly during training using information from the backpropagation. We expand existing work and freely grow neural networks in the form of any Directed Acyclic Graph by reducing expressivity bottlenecks in the architecture. We explore strategies to reduce excessive computations and steer network growth toward more parameter-efficient architectures.
title Growth strategies for arbitrary DAG neural architectures
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2501.12690