WASH: Train your Ensemble with Communication-Efficient Weight Shuffling, then Average

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Fournier, Louis, Nabli, Adel, Aminbeidokhti, Masih, Pedersoli, Marco, Belilovsky, Eugene, Oyallon, Edouard
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917677048528896
author Fournier, Louis
Nabli, Adel
Aminbeidokhti, Masih
Pedersoli, Marco
Belilovsky, Eugene
Oyallon, Edouard
author_facet Fournier, Louis
Nabli, Adel
Aminbeidokhti, Masih
Pedersoli, Marco
Belilovsky, Eugene
Oyallon, Edouard
contents The performance of deep neural networks is enhanced by ensemble methods, which average the output of several models. However, this comes at an increased cost at inference. Weight averaging methods aim at balancing the generalization of ensembling and the inference speed of a single model by averaging the parameters of an ensemble of models. Yet, naive averaging results in poor performance as models converge to different loss basins, and aligning the models to improve the performance of the average is challenging. Alternatively, inspired by distributed training, methods like DART and PAPA have been proposed to train several models in parallel such that they will end up in the same basin, resulting in good averaging accuracy. However, these methods either compromise ensembling accuracy or demand significant communication between models during training. In this paper, we introduce WASH, a novel distributed method for training model ensembles for weight averaging that achieves state-of-the-art image classification accuracy. WASH maintains models within the same basin by randomly shuffling a small percentage of weights during training, resulting in diverse models and lower communication costs compared to standard parameter averaging methods.
format Preprint
id arxiv_https___arxiv_org_abs_2405_17517
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle WASH: Train your Ensemble with Communication-Efficient Weight Shuffling, then Average
Fournier, Louis
Nabli, Adel
Aminbeidokhti, Masih
Pedersoli, Marco
Belilovsky, Eugene
Oyallon, Edouard
Machine Learning
Computer Vision and Pattern Recognition
Neural and Evolutionary Computing
The performance of deep neural networks is enhanced by ensemble methods, which average the output of several models. However, this comes at an increased cost at inference. Weight averaging methods aim at balancing the generalization of ensembling and the inference speed of a single model by averaging the parameters of an ensemble of models. Yet, naive averaging results in poor performance as models converge to different loss basins, and aligning the models to improve the performance of the average is challenging. Alternatively, inspired by distributed training, methods like DART and PAPA have been proposed to train several models in parallel such that they will end up in the same basin, resulting in good averaging accuracy. However, these methods either compromise ensembling accuracy or demand significant communication between models during training. In this paper, we introduce WASH, a novel distributed method for training model ensembles for weight averaging that achieves state-of-the-art image classification accuracy. WASH maintains models within the same basin by randomly shuffling a small percentage of weights during training, resulting in diverse models and lower communication costs compared to standard parameter averaging methods.
title WASH: Train your Ensemble with Communication-Efficient Weight Shuffling, then Average
topic Machine Learning
Computer Vision and Pattern Recognition
Neural and Evolutionary Computing
url https://arxiv.org/abs/2405.17517