Boosting Latent Diffusion with Perceptual Objectives

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Berrada, Tariq, Astolfi, Pietro, Hall, Melissa, Havasi, Marton, Benchetrit, Yohann, Romero-Soriano, Adriana, Alahari, Karteek, Drozdzal, Michal, Verbeek, Jakob
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929682846318592
author Berrada, Tariq
Astolfi, Pietro
Hall, Melissa
Havasi, Marton
Benchetrit, Yohann
Romero-Soriano, Adriana
Alahari, Karteek
Drozdzal, Michal
Verbeek, Jakob
author_facet Berrada, Tariq
Astolfi, Pietro
Hall, Melissa
Havasi, Marton
Benchetrit, Yohann
Romero-Soriano, Adriana
Alahari, Karteek
Drozdzal, Michal
Verbeek, Jakob
contents Latent diffusion models (LDMs) power state-of-the-art high-resolution generative image models. LDMs learn the data distribution in the latent space of an autoencoder (AE) and produce images by mapping the generated latents into RGB image space using the AE decoder. While this approach allows for efficient model training and sampling, it induces a disconnect between the training of the diffusion model and the decoder, resulting in a loss of detail in the generated images. To remediate this disconnect, we propose to leverage the internal features of the decoder to define a latent perceptual loss (LPL). This loss encourages the models to create sharper and more realistic images. Our loss can be seamlessly integrated with common autoencoders used in latent diffusion models, and can be applied to different generative modeling paradigms such as DDPM with epsilon and velocity prediction, as well as flow matching. Extensive experiments with models trained on three datasets at 256 and 512 resolution show improved quantitative -- with boosts between 6% and 20% in FID -- and qualitative results when using our perceptual loss.
format Preprint
id arxiv_https___arxiv_org_abs_2411_04873
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Boosting Latent Diffusion with Perceptual Objectives
Berrada, Tariq
Astolfi, Pietro
Hall, Melissa
Havasi, Marton
Benchetrit, Yohann
Romero-Soriano, Adriana
Alahari, Karteek
Drozdzal, Michal
Verbeek, Jakob
Computer Vision and Pattern Recognition
Latent diffusion models (LDMs) power state-of-the-art high-resolution generative image models. LDMs learn the data distribution in the latent space of an autoencoder (AE) and produce images by mapping the generated latents into RGB image space using the AE decoder. While this approach allows for efficient model training and sampling, it induces a disconnect between the training of the diffusion model and the decoder, resulting in a loss of detail in the generated images. To remediate this disconnect, we propose to leverage the internal features of the decoder to define a latent perceptual loss (LPL). This loss encourages the models to create sharper and more realistic images. Our loss can be seamlessly integrated with common autoencoders used in latent diffusion models, and can be applied to different generative modeling paradigms such as DDPM with epsilon and velocity prediction, as well as flow matching. Extensive experiments with models trained on three datasets at 256 and 512 resolution show improved quantitative -- with boosts between 6% and 20% in FID -- and qualitative results when using our perceptual loss.
title Boosting Latent Diffusion with Perceptual Objectives
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.04873