Divide & Bind Your Attention for Improved Generative Semantic Nursing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Yumeng, Keuper, Margret, Zhang, Dan, Khoreva, Anna
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910525522182144
author Li, Yumeng
Keuper, Margret
Zhang, Dan
Khoreva, Anna
author_facet Li, Yumeng
Keuper, Margret
Zhang, Dan
Khoreva, Anna
contents Emerging large-scale text-to-image generative models, e.g., Stable Diffusion (SD), have exhibited overwhelming results with high fidelity. Despite the magnificent progress, current state-of-the-art models still struggle to generate images fully adhering to the input prompt. Prior work, Attend & Excite, has introduced the concept of Generative Semantic Nursing (GSN), aiming to optimize cross-attention during inference time to better incorporate the semantics. It demonstrates promising results in generating simple prompts, e.g., "a cat and a dog". However, its efficacy declines when dealing with more complex prompts, and it does not explicitly address the problem of improper attribute binding. To address the challenges posed by complex prompts or scenarios involving multiple entities and to achieve improved attribute binding, we propose Divide & Bind. We introduce two novel loss objectives for GSN: a novel attendance loss and a binding loss. Our approach stands out in its ability to faithfully synthesize desired objects with improved attribute alignment from complex prompts and exhibits superior performance across multiple evaluation benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2307_10864
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Divide & Bind Your Attention for Improved Generative Semantic Nursing
Li, Yumeng
Keuper, Margret
Zhang, Dan
Khoreva, Anna
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Emerging large-scale text-to-image generative models, e.g., Stable Diffusion (SD), have exhibited overwhelming results with high fidelity. Despite the magnificent progress, current state-of-the-art models still struggle to generate images fully adhering to the input prompt. Prior work, Attend & Excite, has introduced the concept of Generative Semantic Nursing (GSN), aiming to optimize cross-attention during inference time to better incorporate the semantics. It demonstrates promising results in generating simple prompts, e.g., "a cat and a dog". However, its efficacy declines when dealing with more complex prompts, and it does not explicitly address the problem of improper attribute binding. To address the challenges posed by complex prompts or scenarios involving multiple entities and to achieve improved attribute binding, we propose Divide & Bind. We introduce two novel loss objectives for GSN: a novel attendance loss and a binding loss. Our approach stands out in its ability to faithfully synthesize desired objects with improved attribute alignment from complex prompts and exhibits superior performance across multiple evaluation benchmarks.
title Divide & Bind Your Attention for Improved Generative Semantic Nursing
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2307.10864