Disentanglement with Factor Quantized Variational Autoencoders

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Baykal, Gulcin, Kandemir, Melih, Unal, Gozde
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911249527209984
author Baykal, Gulcin
Kandemir, Melih
Unal, Gozde
author_facet Baykal, Gulcin
Kandemir, Melih
Unal, Gozde
contents Disentangled representation learning aims to represent the underlying generative factors of a dataset in a latent representation independently of one another. In our work, we propose a discrete variational autoencoder (VAE) based model where the ground truth information about the generative factors are not provided to the model. We demonstrate the advantages of learning discrete representations over learning continuous representations in facilitating disentanglement. Furthermore, we propose incorporating an inductive bias into the model to further enhance disentanglement. Precisely, we propose scalar quantization of the latent variables in a latent representation with scalar values from a global codebook, and we add a total correlation term to the optimization as an inductive bias. Our method called FactorQVAE combines optimization based disentanglement approaches with discrete representation learning, and it outperforms the former disentanglement methods in terms of two disentanglement metrics (DCI and InfoMEC) while improving the reconstruction performance. Our code can be found at https://github.com/ituvisionlab/FactorQVAE.
format Preprint
id arxiv_https___arxiv_org_abs_2409_14851
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Disentanglement with Factor Quantized Variational Autoencoders
Baykal, Gulcin
Kandemir, Melih
Unal, Gozde
Computer Vision and Pattern Recognition
Machine Learning
Disentangled representation learning aims to represent the underlying generative factors of a dataset in a latent representation independently of one another. In our work, we propose a discrete variational autoencoder (VAE) based model where the ground truth information about the generative factors are not provided to the model. We demonstrate the advantages of learning discrete representations over learning continuous representations in facilitating disentanglement. Furthermore, we propose incorporating an inductive bias into the model to further enhance disentanglement. Precisely, we propose scalar quantization of the latent variables in a latent representation with scalar values from a global codebook, and we add a total correlation term to the optimization as an inductive bias. Our method called FactorQVAE combines optimization based disentanglement approaches with discrete representation learning, and it outperforms the former disentanglement methods in terms of two disentanglement metrics (DCI and InfoMEC) while improving the reconstruction performance. Our code can be found at https://github.com/ituvisionlab/FactorQVAE.
title Disentanglement with Factor Quantized Variational Autoencoders
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2409.14851