Accelerating Training with Neuron Interaction and Nowcasting Networks

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Knyazev, Boris, Moudgil, Abhinav, Lajoie, Guillaume, Belilovsky, Eugene, Lacoste-Julien, Simon
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910849033043968
author Knyazev, Boris
Moudgil, Abhinav
Lajoie, Guillaume
Belilovsky, Eugene
Lacoste-Julien, Simon
author_facet Knyazev, Boris
Moudgil, Abhinav
Lajoie, Guillaume
Belilovsky, Eugene
Lacoste-Julien, Simon
contents Neural network training can be accelerated when a learnable update rule is used in lieu of classic adaptive optimizers (e.g. Adam). However, learnable update rules can be costly and unstable to train and use. Recently, Jang et al. (2023) proposed a simpler approach to accelerate training based on weight nowcaster networks (WNNs). In their approach, Adam is used for most of the optimization steps and periodically, only every few steps, a WNN nowcasts (predicts near future) parameters. We improve WNNs by proposing neuron interaction and nowcasting (NiNo) networks. In contrast to WNNs, NiNo leverages neuron connectivity and graph neural networks to more accurately nowcast parameters. We further show that in some networks, such as Transformers, modeling neuron connectivity accurately is challenging. We address this and other limitations, which allows NiNo to accelerate Adam training by up to 50% in vision and language tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2409_04434
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Accelerating Training with Neuron Interaction and Nowcasting Networks
Knyazev, Boris
Moudgil, Abhinav
Lajoie, Guillaume
Belilovsky, Eugene
Lacoste-Julien, Simon
Machine Learning
Artificial Intelligence
Neural network training can be accelerated when a learnable update rule is used in lieu of classic adaptive optimizers (e.g. Adam). However, learnable update rules can be costly and unstable to train and use. Recently, Jang et al. (2023) proposed a simpler approach to accelerate training based on weight nowcaster networks (WNNs). In their approach, Adam is used for most of the optimization steps and periodically, only every few steps, a WNN nowcasts (predicts near future) parameters. We improve WNNs by proposing neuron interaction and nowcasting (NiNo) networks. In contrast to WNNs, NiNo leverages neuron connectivity and graph neural networks to more accurately nowcast parameters. We further show that in some networks, such as Transformers, modeling neuron connectivity accurately is challenging. We address this and other limitations, which allows NiNo to accelerate Adam training by up to 50% in vision and language tasks.
title Accelerating Training with Neuron Interaction and Nowcasting Networks
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2409.04434