Training-Free Sketch-Guided Diffusion with Latent Optimization

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ding, Sandra Zhang, Mao, Jiafeng, Aizawa, Kiyoharu
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918012429271040
author Ding, Sandra Zhang
Mao, Jiafeng
Aizawa, Kiyoharu
author_facet Ding, Sandra Zhang
Mao, Jiafeng
Aizawa, Kiyoharu
contents Based on recent advanced diffusion models, Text-to-image (T2I) generation models have demonstrated their capabilities to generate diverse and high-quality images. However, leveraging their potential for real-world content creation, particularly in providing users with precise control over the image generation result, poses a significant challenge. In this paper, we propose an innovative training-free pipeline that extends existing text-to-image generation models to incorporate a sketch as an additional condition. To generate new images with a layout and structure closely resembling the input sketch, we find that these core features of a sketch can be tracked with the cross-attention maps of diffusion models. We introduce latent optimization, a method that refines the noisy latent at each intermediate step of the generation process using cross-attention maps to ensure that the generated images adhere closely to the desired structure outlined in the reference sketch. Through latent optimization, our method enhances the accuracy of image generation, offering users greater control and customization options in content creation.
format Preprint
id arxiv_https___arxiv_org_abs_2409_00313
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Training-Free Sketch-Guided Diffusion with Latent Optimization
Ding, Sandra Zhang
Mao, Jiafeng
Aizawa, Kiyoharu
Computer Vision and Pattern Recognition
Based on recent advanced diffusion models, Text-to-image (T2I) generation models have demonstrated their capabilities to generate diverse and high-quality images. However, leveraging their potential for real-world content creation, particularly in providing users with precise control over the image generation result, poses a significant challenge. In this paper, we propose an innovative training-free pipeline that extends existing text-to-image generation models to incorporate a sketch as an additional condition. To generate new images with a layout and structure closely resembling the input sketch, we find that these core features of a sketch can be tracked with the cross-attention maps of diffusion models. We introduce latent optimization, a method that refines the noisy latent at each intermediate step of the generation process using cross-attention maps to ensure that the generated images adhere closely to the desired structure outlined in the reference sketch. Through latent optimization, our method enhances the accuracy of image generation, offering users greater control and customization options in content creation.
title Training-Free Sketch-Guided Diffusion with Latent Optimization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2409.00313