Enhancing Diversity in Bayesian Deep Learning via Hyperspherical Energy Minimization of CKA

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Smerkous, David, Bai, Qinxun, Li, Fuxin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910680418877440
author Smerkous, David
Bai, Qinxun
Li, Fuxin
author_facet Smerkous, David
Bai, Qinxun
Li, Fuxin
contents Particle-based Bayesian deep learning often requires a similarity metric to compare two networks. However, naive similarity metrics lack permutation invariance and are inappropriate for comparing networks. Centered Kernel Alignment (CKA) on feature kernels has been proposed to compare deep networks but has not been used as an optimization objective in Bayesian deep learning. In this paper, we explore the use of CKA in Bayesian deep learning to generate diverse ensembles and hypernetworks that output a network posterior. Noting that CKA projects kernels onto a unit hypersphere and that directly optimizing the CKA objective leads to diminishing gradients when two networks are very similar. We propose adopting the approach of hyperspherical energy (HE) on top of CKA kernels to address this drawback and improve training stability. Additionally, by leveraging CKA-based feature kernels, we derive feature repulsive terms applied to synthetically generated outlier examples. Experiments on both diverse ensembles and hypernetworks show that our approach significantly outperforms baselines in terms of uncertainty quantification in both synthetic and realistic outlier detection tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2411_00259
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Diversity in Bayesian Deep Learning via Hyperspherical Energy Minimization of CKA
Smerkous, David
Bai, Qinxun
Li, Fuxin
Machine Learning
Particle-based Bayesian deep learning often requires a similarity metric to compare two networks. However, naive similarity metrics lack permutation invariance and are inappropriate for comparing networks. Centered Kernel Alignment (CKA) on feature kernels has been proposed to compare deep networks but has not been used as an optimization objective in Bayesian deep learning. In this paper, we explore the use of CKA in Bayesian deep learning to generate diverse ensembles and hypernetworks that output a network posterior. Noting that CKA projects kernels onto a unit hypersphere and that directly optimizing the CKA objective leads to diminishing gradients when two networks are very similar. We propose adopting the approach of hyperspherical energy (HE) on top of CKA kernels to address this drawback and improve training stability. Additionally, by leveraging CKA-based feature kernels, we derive feature repulsive terms applied to synthetically generated outlier examples. Experiments on both diverse ensembles and hypernetworks show that our approach significantly outperforms baselines in terms of uncertainty quantification in both synthetic and realistic outlier detection tasks.
title Enhancing Diversity in Bayesian Deep Learning via Hyperspherical Energy Minimization of CKA
topic Machine Learning
url https://arxiv.org/abs/2411.00259