DC-AE 1.5: Accelerating Diffusion Model Convergence with Structured Latent Space

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Junyu, Zou, Dongyun, He, Wenkun, Chen, Junsong, Xie, Enze, Han, Song, Cai, Han
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916875110187008
author Chen, Junyu
Zou, Dongyun
He, Wenkun
Chen, Junsong
Xie, Enze
Han, Song
Cai, Han
author_facet Chen, Junyu
Zou, Dongyun
He, Wenkun
Chen, Junsong
Xie, Enze
Han, Song
Cai, Han
contents We present DC-AE 1.5, a new family of deep compression autoencoders for high-resolution diffusion models. Increasing the autoencoder's latent channel number is a highly effective approach for improving its reconstruction quality. However, it results in slow convergence for diffusion models, leading to poorer generation quality despite better reconstruction quality. This issue limits the quality upper bound of latent diffusion models and hinders the employment of autoencoders with higher spatial compression ratios. We introduce two key innovations to address this challenge: i) Structured Latent Space, a training-based approach to impose a desired channel-wise structure on the latent space with front latent channels capturing object structures and latter latent channels capturing image details; ii) Augmented Diffusion Training, an augmented diffusion training strategy with additional diffusion training objectives on object latent channels to accelerate convergence. With these techniques, DC-AE 1.5 delivers faster convergence and better diffusion scaling results than DC-AE. On ImageNet 512x512, DC-AE-1.5-f64c128 delivers better image generation quality than DC-AE-f32c32 while being 4x faster. Code: https://github.com/dc-ai-projects/DC-Gen.
format Preprint
id arxiv_https___arxiv_org_abs_2508_00413
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DC-AE 1.5: Accelerating Diffusion Model Convergence with Structured Latent Space
Chen, Junyu
Zou, Dongyun
He, Wenkun
Chen, Junsong
Xie, Enze
Han, Song
Cai, Han
Computer Vision and Pattern Recognition
Artificial Intelligence
We present DC-AE 1.5, a new family of deep compression autoencoders for high-resolution diffusion models. Increasing the autoencoder's latent channel number is a highly effective approach for improving its reconstruction quality. However, it results in slow convergence for diffusion models, leading to poorer generation quality despite better reconstruction quality. This issue limits the quality upper bound of latent diffusion models and hinders the employment of autoencoders with higher spatial compression ratios. We introduce two key innovations to address this challenge: i) Structured Latent Space, a training-based approach to impose a desired channel-wise structure on the latent space with front latent channels capturing object structures and latter latent channels capturing image details; ii) Augmented Diffusion Training, an augmented diffusion training strategy with additional diffusion training objectives on object latent channels to accelerate convergence. With these techniques, DC-AE 1.5 delivers faster convergence and better diffusion scaling results than DC-AE. On ImageNet 512x512, DC-AE-1.5-f64c128 delivers better image generation quality than DC-AE-f32c32 while being 4x faster. Code: https://github.com/dc-ai-projects/DC-Gen.
title DC-AE 1.5: Accelerating Diffusion Model Convergence with Structured Latent Space
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2508.00413