Diffusion-Based Conditional Image Editing through Optimized Inference with Guidance

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lee, Hyunsoo, Kang, Minsoo, Han, Bohyung
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913621267709952
author Lee, Hyunsoo
Kang, Minsoo
Han, Bohyung
author_facet Lee, Hyunsoo
Kang, Minsoo
Han, Bohyung
contents We present a simple but effective training-free approach for text-driven image-to-image translation based on a pretrained text-to-image diffusion model. Our goal is to generate an image that aligns with the target task while preserving the structure and background of a source image. To this end, we derive the representation guidance with a combination of two objectives: maximizing the similarity to the target prompt based on the CLIP score and minimizing the structural distance to the source latent variable. This guidance improves the fidelity of the generated target image to the given target prompt while maintaining the structure integrity of the source image. To incorporate the representation guidance component, we optimize the target latent variable of diffusion model's reverse process with the guidance. Experimental results demonstrate that our method achieves outstanding image-to-image translation performance on various tasks when combined with the pretrained Stable Diffusion model.
format Preprint
id arxiv_https___arxiv_org_abs_2412_15798
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Diffusion-Based Conditional Image Editing through Optimized Inference with Guidance
Lee, Hyunsoo
Kang, Minsoo
Han, Bohyung
Computer Vision and Pattern Recognition
We present a simple but effective training-free approach for text-driven image-to-image translation based on a pretrained text-to-image diffusion model. Our goal is to generate an image that aligns with the target task while preserving the structure and background of a source image. To this end, we derive the representation guidance with a combination of two objectives: maximizing the similarity to the target prompt based on the CLIP score and minimizing the structural distance to the source latent variable. This guidance improves the fidelity of the generated target image to the given target prompt while maintaining the structure integrity of the source image. To incorporate the representation guidance component, we optimize the target latent variable of diffusion model's reverse process with the guidance. Experimental results demonstrate that our method achieves outstanding image-to-image translation performance on various tasks when combined with the pretrained Stable Diffusion model.
title Diffusion-Based Conditional Image Editing through Optimized Inference with Guidance
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.15798