Sequential Bayesian Neural Subnetwork Ensembles

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jantre, Sanket, Bhattacharya, Shrijita, Urban, Nathan M., Yoon, Byung-Jun, Maiti, Tapabrata, Balaprakash, Prasanna, Madireddy, Sandeep
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911994699841536
author Jantre, Sanket
Bhattacharya, Shrijita
Urban, Nathan M.
Yoon, Byung-Jun
Maiti, Tapabrata
Balaprakash, Prasanna
Madireddy, Sandeep
author_facet Jantre, Sanket
Bhattacharya, Shrijita
Urban, Nathan M.
Yoon, Byung-Jun
Maiti, Tapabrata
Balaprakash, Prasanna
Madireddy, Sandeep
contents Deep ensembles have emerged as a powerful technique for improving predictive performance and enhancing model robustness across various applications by leveraging model diversity. However, traditional deep ensemble methods are often computationally expensive and rely on deterministic models, which may limit their flexibility. Additionally, while sparse subnetworks of dense models have shown promise in matching the performance of their dense counterparts and even enhancing robustness, existing methods for inducing sparsity typically incur training costs comparable to those of training a single dense model, as they either gradually prune the network during training or apply thresholding post-training. In light of these challenges, we propose an approach for sequential ensembling of dynamic Bayesian neural subnetworks that consistently maintains reduced model complexity throughout the training process while generating diverse ensembles in a single forward pass. Our approach involves an initial exploration phase to identify high-performing regions within the parameter space, followed by multiple exploitation phases that take advantage of the compactness of the sparse model. These exploitation phases quickly converge to different minima in the energy landscape, corresponding to high-performing subnetworks that together form a diverse and robust ensemble. We empirically demonstrate that our proposed approach outperforms traditional dense and sparse deterministic and Bayesian ensemble models in terms of prediction accuracy, uncertainty estimation, out-of-distribution detection, and adversarial robustness.
format Preprint
id arxiv_https___arxiv_org_abs_2206_00794
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Sequential Bayesian Neural Subnetwork Ensembles
Jantre, Sanket
Bhattacharya, Shrijita
Urban, Nathan M.
Yoon, Byung-Jun
Maiti, Tapabrata
Balaprakash, Prasanna
Madireddy, Sandeep
Machine Learning
Statistics Theory
Deep ensembles have emerged as a powerful technique for improving predictive performance and enhancing model robustness across various applications by leveraging model diversity. However, traditional deep ensemble methods are often computationally expensive and rely on deterministic models, which may limit their flexibility. Additionally, while sparse subnetworks of dense models have shown promise in matching the performance of their dense counterparts and even enhancing robustness, existing methods for inducing sparsity typically incur training costs comparable to those of training a single dense model, as they either gradually prune the network during training or apply thresholding post-training. In light of these challenges, we propose an approach for sequential ensembling of dynamic Bayesian neural subnetworks that consistently maintains reduced model complexity throughout the training process while generating diverse ensembles in a single forward pass. Our approach involves an initial exploration phase to identify high-performing regions within the parameter space, followed by multiple exploitation phases that take advantage of the compactness of the sparse model. These exploitation phases quickly converge to different minima in the energy landscape, corresponding to high-performing subnetworks that together form a diverse and robust ensemble. We empirically demonstrate that our proposed approach outperforms traditional dense and sparse deterministic and Bayesian ensemble models in terms of prediction accuracy, uncertainty estimation, out-of-distribution detection, and adversarial robustness.
title Sequential Bayesian Neural Subnetwork Ensembles
topic Machine Learning
Statistics Theory
url https://arxiv.org/abs/2206.00794