Latent Diffusion Model without Variational Autoencoder

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shi, Minglei, Wang, Haolin, Zheng, Wenzhao, Yuan, Ziyang, Wu, Xiaoshi, Wang, Xintao, Wan, Pengfei, Zhou, Jie, Lu, Jiwen
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911477614510080
author Shi, Minglei
Wang, Haolin
Zheng, Wenzhao
Yuan, Ziyang
Wu, Xiaoshi
Wang, Xintao
Wan, Pengfei
Zhou, Jie
Lu, Jiwen
author_facet Shi, Minglei
Wang, Haolin
Zheng, Wenzhao
Yuan, Ziyang
Wu, Xiaoshi
Wang, Xintao
Wan, Pengfei
Zhou, Jie
Lu, Jiwen
contents Recent progress in diffusion-based visual generation has largely relied on latent diffusion models with variational autoencoders (VAEs). While effective for high-fidelity synthesis, this VAE+diffusion paradigm suffers from limited training efficiency, slow inference, and poor transferability to broader vision tasks. These issues stem from a key limitation of VAE latent spaces: the lack of clear semantic separation and strong discriminative structure. Our analysis confirms that these properties are crucial not only for perception and understanding tasks, but also for the stable and efficient training of latent diffusion models. Motivated by this insight, we introduce SVG, a novel latent diffusion model without variational autoencoders, which leverages self-supervised representations for visual generation. SVG constructs a feature space with clear semantic discriminability by leveraging frozen DINO features, while a lightweight residual branch captures fine-grained details for high-fidelity reconstruction. Diffusion models are trained directly on this semantically structured latent space to facilitate more efficient learning. As a result, SVG enables accelerated diffusion training, supports few-step sampling, and improves generative quality. Experimental results further show that SVG preserves the semantic and discriminative capabilities of the underlying self-supervised representations, providing a principled pathway toward task-general, high-quality visual representations. Code and interpretations are available at https://howlin-wang.github.io/svg/.
format Preprint
id arxiv_https___arxiv_org_abs_2510_15301
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Latent Diffusion Model without Variational Autoencoder
Shi, Minglei
Wang, Haolin
Zheng, Wenzhao
Yuan, Ziyang
Wu, Xiaoshi
Wang, Xintao
Wan, Pengfei
Zhou, Jie
Lu, Jiwen
Computer Vision and Pattern Recognition
Artificial Intelligence
Recent progress in diffusion-based visual generation has largely relied on latent diffusion models with variational autoencoders (VAEs). While effective for high-fidelity synthesis, this VAE+diffusion paradigm suffers from limited training efficiency, slow inference, and poor transferability to broader vision tasks. These issues stem from a key limitation of VAE latent spaces: the lack of clear semantic separation and strong discriminative structure. Our analysis confirms that these properties are crucial not only for perception and understanding tasks, but also for the stable and efficient training of latent diffusion models. Motivated by this insight, we introduce SVG, a novel latent diffusion model without variational autoencoders, which leverages self-supervised representations for visual generation. SVG constructs a feature space with clear semantic discriminability by leveraging frozen DINO features, while a lightweight residual branch captures fine-grained details for high-fidelity reconstruction. Diffusion models are trained directly on this semantically structured latent space to facilitate more efficient learning. As a result, SVG enables accelerated diffusion training, supports few-step sampling, and improves generative quality. Experimental results further show that SVG preserves the semantic and discriminative capabilities of the underlying self-supervised representations, providing a principled pathway toward task-general, high-quality visual representations. Code and interpretations are available at https://howlin-wang.github.io/svg/.
title Latent Diffusion Model without Variational Autoencoder
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2510.15301