Diffusion Bridge AutoEncoders for Unsupervised Representation Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Yeongmin, Lee, Kwanghyeon, Park, Minsang, Na, Byeonghu, Moon, Il-Chul
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915186803212288
author Kim, Yeongmin
Lee, Kwanghyeon
Park, Minsang
Na, Byeonghu
Moon, Il-Chul
author_facet Kim, Yeongmin
Lee, Kwanghyeon
Park, Minsang
Na, Byeonghu
Moon, Il-Chul
contents Diffusion-based representation learning has achieved substantial attention due to its promising capabilities in latent representation and sample generation. Recent studies have employed an auxiliary encoder to identify a corresponding representation from a sample and to adjust the dimensionality of a latent variable z. Meanwhile, this auxiliary structure invokes information split problem because the diffusion and the auxiliary encoder would divide the information from the sample into two representations for each model. Particularly, the information modeled by the diffusion becomes over-regularized because of the static prior distribution on xT. To address this problem, we introduce Diffusion Bridge AuteEncoders (DBAE), which enable z-dependent endpoint xT inference through a feed-forward architecture. This structure creates an information bottleneck at z, so xT becomes dependent on z in its generation. This results in two consequences: 1) z holds the full information of samples, and 2) xT becomes a learnable distribution, not static any further. We propose an objective function for DBAE to enable both reconstruction and generative modeling, with their theoretical justification. Empirical evidence supports the effectiveness of the intended design in DBAE, which notably enhances downstream inference quality, reconstruction, and disentanglement. Additionally, DBAE generates high-fidelity samples in the unconditional generation. Our code is available at https://github.com/aailab-kaist/DBAE.
format Preprint
id arxiv_https___arxiv_org_abs_2405_17111
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Diffusion Bridge AutoEncoders for Unsupervised Representation Learning
Kim, Yeongmin
Lee, Kwanghyeon
Park, Minsang
Na, Byeonghu
Moon, Il-Chul
Machine Learning
Diffusion-based representation learning has achieved substantial attention due to its promising capabilities in latent representation and sample generation. Recent studies have employed an auxiliary encoder to identify a corresponding representation from a sample and to adjust the dimensionality of a latent variable z. Meanwhile, this auxiliary structure invokes information split problem because the diffusion and the auxiliary encoder would divide the information from the sample into two representations for each model. Particularly, the information modeled by the diffusion becomes over-regularized because of the static prior distribution on xT. To address this problem, we introduce Diffusion Bridge AuteEncoders (DBAE), which enable z-dependent endpoint xT inference through a feed-forward architecture. This structure creates an information bottleneck at z, so xT becomes dependent on z in its generation. This results in two consequences: 1) z holds the full information of samples, and 2) xT becomes a learnable distribution, not static any further. We propose an objective function for DBAE to enable both reconstruction and generative modeling, with their theoretical justification. Empirical evidence supports the effectiveness of the intended design in DBAE, which notably enhances downstream inference quality, reconstruction, and disentanglement. Additionally, DBAE generates high-fidelity samples in the unconditional generation. Our code is available at https://github.com/aailab-kaist/DBAE.
title Diffusion Bridge AutoEncoders for Unsupervised Representation Learning
topic Machine Learning
url https://arxiv.org/abs/2405.17111