EdVAE: Mitigating Codebook Collapse with Evidential Discrete Variational Autoencoders

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Baykal, Gulcin, Kandemir, Melih, Unal, Gozde
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909254666944512
author Baykal, Gulcin
Kandemir, Melih
Unal, Gozde
author_facet Baykal, Gulcin
Kandemir, Melih
Unal, Gozde
contents Codebook collapse is a common problem in training deep generative models with discrete representation spaces like Vector Quantized Variational Autoencoders (VQ-VAEs). We observe that the same problem arises for the alternatively designed discrete variational autoencoders (dVAEs) whose encoder directly learns a distribution over the codebook embeddings to represent the data. We hypothesize that using the softmax function to obtain a probability distribution causes the codebook collapse by assigning overconfident probabilities to the best matching codebook elements. In this paper, we propose a novel way to incorporate evidential deep learning (EDL) instead of softmax to combat the codebook collapse problem of dVAE. We evidentially monitor the significance of attaining the probability distribution over the codebook embeddings, in contrast to softmax usage. Our experiments using various datasets show that our model, called EdVAE, mitigates codebook collapse while improving the reconstruction performance, and enhances the codebook usage compared to dVAE and VQ-VAE based models. Our code can be found at https://github.com/ituvisionlab/EdVAE .
format Preprint
id arxiv_https___arxiv_org_abs_2310_05718
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle EdVAE: Mitigating Codebook Collapse with Evidential Discrete Variational Autoencoders
Baykal, Gulcin
Kandemir, Melih
Unal, Gozde
Computer Vision and Pattern Recognition
Machine Learning
Codebook collapse is a common problem in training deep generative models with discrete representation spaces like Vector Quantized Variational Autoencoders (VQ-VAEs). We observe that the same problem arises for the alternatively designed discrete variational autoencoders (dVAEs) whose encoder directly learns a distribution over the codebook embeddings to represent the data. We hypothesize that using the softmax function to obtain a probability distribution causes the codebook collapse by assigning overconfident probabilities to the best matching codebook elements. In this paper, we propose a novel way to incorporate evidential deep learning (EDL) instead of softmax to combat the codebook collapse problem of dVAE. We evidentially monitor the significance of attaining the probability distribution over the codebook embeddings, in contrast to softmax usage. Our experiments using various datasets show that our model, called EdVAE, mitigates codebook collapse while improving the reconstruction performance, and enhances the codebook usage compared to dVAE and VQ-VAE based models. Our code can be found at https://github.com/ituvisionlab/EdVAE .
title EdVAE: Mitigating Codebook Collapse with Evidential Discrete Variational Autoencoders
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2310.05718