VSC: Visual Search Compositional Text-to-Image Diffusion Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dat, Do Huu, Hyeonu, Nam, Mao, Po-Yuan, Oh, Tae-Hyun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912357765087232
author Dat, Do Huu
Hyeonu, Nam
Mao, Po-Yuan
Oh, Tae-Hyun
author_facet Dat, Do Huu
Hyeonu, Nam
Mao, Po-Yuan
Oh, Tae-Hyun
contents Text-to-image diffusion models have shown impressive capabilities in generating realistic visuals from natural-language prompts, yet they often struggle with accurately binding attributes to corresponding objects, especially in prompts containing multiple attribute-object pairs. This challenge primarily arises from the limitations of commonly used text encoders, such as CLIP, which can fail to encode complex linguistic relationships and modifiers effectively. Existing approaches have attempted to mitigate these issues through attention map control during inference and the use of layout information or fine-tuning during training, yet they face performance drops with increased prompt complexity. In this work, we introduce a novel compositional generation method that leverages pairwise image embeddings to improve attribute-object binding. Our approach decomposes complex prompts into sub-prompts, generates corresponding images, and computes visual prototypes that fuse with text embeddings to enhance representation. By applying segmentation-based localization training, we address cross-attention misalignment, achieving improved accuracy in binding multiple attributes to objects. Our approaches outperform existing compositional text-to-image diffusion models on the benchmark T2I CompBench, achieving better image quality, evaluated by humans, and emerging robustness under scaling number of binding pairs in the prompt.
format Preprint
id arxiv_https___arxiv_org_abs_2505_01104
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VSC: Visual Search Compositional Text-to-Image Diffusion Model
Dat, Do Huu
Hyeonu, Nam
Mao, Po-Yuan
Oh, Tae-Hyun
Computer Vision and Pattern Recognition
Text-to-image diffusion models have shown impressive capabilities in generating realistic visuals from natural-language prompts, yet they often struggle with accurately binding attributes to corresponding objects, especially in prompts containing multiple attribute-object pairs. This challenge primarily arises from the limitations of commonly used text encoders, such as CLIP, which can fail to encode complex linguistic relationships and modifiers effectively. Existing approaches have attempted to mitigate these issues through attention map control during inference and the use of layout information or fine-tuning during training, yet they face performance drops with increased prompt complexity. In this work, we introduce a novel compositional generation method that leverages pairwise image embeddings to improve attribute-object binding. Our approach decomposes complex prompts into sub-prompts, generates corresponding images, and computes visual prototypes that fuse with text embeddings to enhance representation. By applying segmentation-based localization training, we address cross-attention misalignment, achieving improved accuracy in binding multiple attributes to objects. Our approaches outperform existing compositional text-to-image diffusion models on the benchmark T2I CompBench, achieving better image quality, evaluated by humans, and emerging robustness under scaling number of binding pairs in the prompt.
title VSC: Visual Search Compositional Text-to-Image Diffusion Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.01104