Enhancing Image Layout Control with Loss-Guided Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Patel, Zakaria, Serkh, Kirill
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917777519935488
author Patel, Zakaria
Serkh, Kirill
author_facet Patel, Zakaria
Serkh, Kirill
contents Diffusion models are a powerful class of generative models capable of producing high-quality images from pure noise using a simple text prompt. While most methods which introduce additional spatial constraints into the generated images (e.g., bounding boxes) require fine-tuning, a smaller and more recent subset of these methods take advantage of the models' attention mechanism, and are training-free. These methods generally fall into one of two categories. The first entails modifying the cross-attention maps of specific tokens directly to enhance the signal in certain regions of the image. The second works by defining a loss function over the cross-attention maps, and using the gradient of this loss to guide the latent. While previous work explores these as alternative strategies, we provide an interpretation for these methods which highlights their complimentary features, and demonstrate that it is possible to obtain superior performance when both methods are used in concert.
format Preprint
id arxiv_https___arxiv_org_abs_2405_14101
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Image Layout Control with Loss-Guided Diffusion Models
Patel, Zakaria
Serkh, Kirill
Computer Vision and Pattern Recognition
Graphics
Machine Learning
Diffusion models are a powerful class of generative models capable of producing high-quality images from pure noise using a simple text prompt. While most methods which introduce additional spatial constraints into the generated images (e.g., bounding boxes) require fine-tuning, a smaller and more recent subset of these methods take advantage of the models' attention mechanism, and are training-free. These methods generally fall into one of two categories. The first entails modifying the cross-attention maps of specific tokens directly to enhance the signal in certain regions of the image. The second works by defining a loss function over the cross-attention maps, and using the gradient of this loss to guide the latent. While previous work explores these as alternative strategies, we provide an interpretation for these methods which highlights their complimentary features, and demonstrate that it is possible to obtain superior performance when both methods are used in concert.
title Enhancing Image Layout Control with Loss-Guided Diffusion Models
topic Computer Vision and Pattern Recognition
Graphics
Machine Learning
url https://arxiv.org/abs/2405.14101