DGAE: Diffusion-Guided Autoencoder for Efficient Latent Representation Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Dongxu, Zhu, Jiahui, Peng, Yuang, Tang, Haomiao, Chen, Yuwei, Han, Chunrui, Ge, Zheng, Jiang, Daxin, Liao, Mingxue
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911373293780992
author Liu, Dongxu
Zhu, Jiahui
Peng, Yuang
Tang, Haomiao
Chen, Yuwei
Han, Chunrui
Ge, Zheng
Jiang, Daxin
Liao, Mingxue
author_facet Liu, Dongxu
Zhu, Jiahui
Peng, Yuang
Tang, Haomiao
Chen, Yuwei
Han, Chunrui
Ge, Zheng
Jiang, Daxin
Liao, Mingxue
contents Autoencoders empower state-of-the-art image and video generative models by compressing pixels into a latent space through visual tokenization. Although recent advances have alleviated the performance degradation of autoencoders under high compression ratios, addressing the training instability caused by GAN remains an open challenge. While improving spatial compression, we also aim to minimize the latent space dimensionality, enabling more efficient and compact representations. To tackle these challenges, we focus on improving the decoder's expressiveness. Concretely, we propose DGAE, which employs a diffusion model to guide the decoder in recovering informative signals that are not fully decoded from the latent representation. With this design, DGAE effectively mitigates the performance degradation under high spatial compression rates. At the same time, DGAE achieves state-of-the-art performance with a 2x smaller latent space. When integrated with Diffusion Models, DGAE demonstrates competitive performance on image generation for ImageNet-1K and shows that this compact latent representation facilitates faster convergence of the diffusion model.
format Preprint
id arxiv_https___arxiv_org_abs_2506_09644
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DGAE: Diffusion-Guided Autoencoder for Efficient Latent Representation Learning
Liu, Dongxu
Zhu, Jiahui
Peng, Yuang
Tang, Haomiao
Chen, Yuwei
Han, Chunrui
Ge, Zheng
Jiang, Daxin
Liao, Mingxue
Computer Vision and Pattern Recognition
Artificial Intelligence
Autoencoders empower state-of-the-art image and video generative models by compressing pixels into a latent space through visual tokenization. Although recent advances have alleviated the performance degradation of autoencoders under high compression ratios, addressing the training instability caused by GAN remains an open challenge. While improving spatial compression, we also aim to minimize the latent space dimensionality, enabling more efficient and compact representations. To tackle these challenges, we focus on improving the decoder's expressiveness. Concretely, we propose DGAE, which employs a diffusion model to guide the decoder in recovering informative signals that are not fully decoded from the latent representation. With this design, DGAE effectively mitigates the performance degradation under high spatial compression rates. At the same time, DGAE achieves state-of-the-art performance with a 2x smaller latent space. When integrated with Diffusion Models, DGAE demonstrates competitive performance on image generation for ImageNet-1K and shows that this compact latent representation facilitates faster convergence of the diffusion model.
title DGAE: Diffusion-Guided Autoencoder for Efficient Latent Representation Learning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2506.09644