Be Yourself: Bounded Attention for Multi-Subject Text-to-Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dahary, Omer, Patashnik, Or, Aberman, Kfir, Cohen-Or, Daniel
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916175711043584
author Dahary, Omer
Patashnik, Or
Aberman, Kfir
Cohen-Or, Daniel
author_facet Dahary, Omer
Patashnik, Or
Aberman, Kfir
Cohen-Or, Daniel
contents Text-to-image diffusion models have an unprecedented ability to generate diverse and high-quality images. However, they often struggle to faithfully capture the intended semantics of complex input prompts that include multiple subjects. Recently, numerous layout-to-image extensions have been introduced to improve user control, aiming to localize subjects represented by specific tokens. Yet, these methods often produce semantically inaccurate images, especially when dealing with multiple semantically or visually similar subjects. In this work, we study and analyze the causes of these limitations. Our exploration reveals that the primary issue stems from inadvertent semantic leakage between subjects in the denoising process. This leakage is attributed to the diffusion model's attention layers, which tend to blend the visual features of different subjects. To address these issues, we introduce Bounded Attention, a training-free method for bounding the information flow in the sampling process. Bounded Attention prevents detrimental leakage among subjects and enables guiding the generation to promote each subject's individuality, even with complex multi-subject conditioning. Through extensive experimentation, we demonstrate that our method empowers the generation of multiple subjects that better align with given prompts and layouts.
format Preprint
id arxiv_https___arxiv_org_abs_2403_16990
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Be Yourself: Bounded Attention for Multi-Subject Text-to-Image Generation
Dahary, Omer
Patashnik, Or
Aberman, Kfir
Cohen-Or, Daniel
Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Machine Learning
Text-to-image diffusion models have an unprecedented ability to generate diverse and high-quality images. However, they often struggle to faithfully capture the intended semantics of complex input prompts that include multiple subjects. Recently, numerous layout-to-image extensions have been introduced to improve user control, aiming to localize subjects represented by specific tokens. Yet, these methods often produce semantically inaccurate images, especially when dealing with multiple semantically or visually similar subjects. In this work, we study and analyze the causes of these limitations. Our exploration reveals that the primary issue stems from inadvertent semantic leakage between subjects in the denoising process. This leakage is attributed to the diffusion model's attention layers, which tend to blend the visual features of different subjects. To address these issues, we introduce Bounded Attention, a training-free method for bounding the information flow in the sampling process. Bounded Attention prevents detrimental leakage among subjects and enables guiding the generation to promote each subject's individuality, even with complex multi-subject conditioning. Through extensive experimentation, we demonstrate that our method empowers the generation of multiple subjects that better align with given prompts and layouts.
title Be Yourself: Bounded Attention for Multi-Subject Text-to-Image Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Machine Learning
url https://arxiv.org/abs/2403.16990