Accelerating Training with Neuron Interaction and Nowcasting Networks
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866910849033043968 |
|---|---|
| author | Knyazev, Boris Moudgil, Abhinav Lajoie, Guillaume Belilovsky, Eugene Lacoste-Julien, Simon |
| author_facet | Knyazev, Boris Moudgil, Abhinav Lajoie, Guillaume Belilovsky, Eugene Lacoste-Julien, Simon |
| contents | Neural network training can be accelerated when a learnable update rule is used in lieu of classic adaptive optimizers (e.g. Adam). However, learnable update rules can be costly and unstable to train and use. Recently, Jang et al. (2023) proposed a simpler approach to accelerate training based on weight nowcaster networks (WNNs). In their approach, Adam is used for most of the optimization steps and periodically, only every few steps, a WNN nowcasts (predicts near future) parameters. We improve WNNs by proposing neuron interaction and nowcasting (NiNo) networks. In contrast to WNNs, NiNo leverages neuron connectivity and graph neural networks to more accurately nowcast parameters. We further show that in some networks, such as Transformers, modeling neuron connectivity accurately is challenging. We address this and other limitations, which allows NiNo to accelerate Adam training by up to 50% in vision and language tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2409_04434 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Accelerating Training with Neuron Interaction and Nowcasting Networks Knyazev, Boris Moudgil, Abhinav Lajoie, Guillaume Belilovsky, Eugene Lacoste-Julien, Simon Machine Learning Artificial Intelligence Neural network training can be accelerated when a learnable update rule is used in lieu of classic adaptive optimizers (e.g. Adam). However, learnable update rules can be costly and unstable to train and use. Recently, Jang et al. (2023) proposed a simpler approach to accelerate training based on weight nowcaster networks (WNNs). In their approach, Adam is used for most of the optimization steps and periodically, only every few steps, a WNN nowcasts (predicts near future) parameters. We improve WNNs by proposing neuron interaction and nowcasting (NiNo) networks. In contrast to WNNs, NiNo leverages neuron connectivity and graph neural networks to more accurately nowcast parameters. We further show that in some networks, such as Transformers, modeling neuron connectivity accurately is challenging. We address this and other limitations, which allows NiNo to accelerate Adam training by up to 50% in vision and language tasks. |
| title | Accelerating Training with Neuron Interaction and Nowcasting Networks |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2409.04434 |