Fine-Grained Alignment and Noise Refinement for Compositional Text-to-Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Izadi, Amir Mohammad, Hosseini, Seyed Mohammad Hadi, Tabar, Soroush Vafaie, Abdollahi, Ali, Saghafian, Armin, Baghshah, Mahdieh Soleymani
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916857698582528
author Izadi, Amir Mohammad
Hosseini, Seyed Mohammad Hadi
Tabar, Soroush Vafaie
Abdollahi, Ali
Saghafian, Armin
Baghshah, Mahdieh Soleymani
author_facet Izadi, Amir Mohammad
Hosseini, Seyed Mohammad Hadi
Tabar, Soroush Vafaie
Abdollahi, Ali
Saghafian, Armin
Baghshah, Mahdieh Soleymani
contents Text-to-image generative models have made significant advancements in recent years; however, accurately capturing intricate details in textual prompts-such as entity missing, attribute binding errors, and incorrect relationships remains a formidable challenge. In response, we present an innovative, training-free method that directly addresses these challenges by incorporating tailored objectives to account for textual constraints. Unlike layout-based approaches that enforce rigid structures and limit diversity, our proposed approach offers a more flexible arrangement of the scene by imposing just the extracted constraints from the text, without any unnecessary additions. These constraints are formulated as losses-entity missing, entity mixing, attribute binding, and spatial relationships-integrated into a unified loss that is applied in the first generation stage. Furthermore, we introduce a feedback-driven system for fine-grained initial noise refinement. This system integrates a verifier that evaluates the generated image, identifies inconsistencies, and provides corrective feedback. Leveraging this feedback, our refinement method first targets the unmet constraints by refining the faulty attention maps caused by initial noise, through the optimization of selective losses associated with these constraints. Subsequently, our unified loss function is reapplied to proceed the second generation phase. Experimental results demonstrate that our method, relying solely on our proposed objective functions, significantly enhances compositionality, achieving a 24% improvement in human evaluation and a 25% gain in spatial relationships. Furthermore, our fine-grained noise refinement proves effective, boosting performance by up to 5%. Code is available at \href{https://github.com/hadi-hosseini/noise-refinement}{https://github.com/hadi-hosseini/noise-refinement}.
format Preprint
id arxiv_https___arxiv_org_abs_2503_06506
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fine-Grained Alignment and Noise Refinement for Compositional Text-to-Image Generation
Izadi, Amir Mohammad
Hosseini, Seyed Mohammad Hadi
Tabar, Soroush Vafaie
Abdollahi, Ali
Saghafian, Armin
Baghshah, Mahdieh Soleymani
Computer Vision and Pattern Recognition
Machine Learning
Text-to-image generative models have made significant advancements in recent years; however, accurately capturing intricate details in textual prompts-such as entity missing, attribute binding errors, and incorrect relationships remains a formidable challenge. In response, we present an innovative, training-free method that directly addresses these challenges by incorporating tailored objectives to account for textual constraints. Unlike layout-based approaches that enforce rigid structures and limit diversity, our proposed approach offers a more flexible arrangement of the scene by imposing just the extracted constraints from the text, without any unnecessary additions. These constraints are formulated as losses-entity missing, entity mixing, attribute binding, and spatial relationships-integrated into a unified loss that is applied in the first generation stage. Furthermore, we introduce a feedback-driven system for fine-grained initial noise refinement. This system integrates a verifier that evaluates the generated image, identifies inconsistencies, and provides corrective feedback. Leveraging this feedback, our refinement method first targets the unmet constraints by refining the faulty attention maps caused by initial noise, through the optimization of selective losses associated with these constraints. Subsequently, our unified loss function is reapplied to proceed the second generation phase. Experimental results demonstrate that our method, relying solely on our proposed objective functions, significantly enhances compositionality, achieving a 24% improvement in human evaluation and a 25% gain in spatial relationships. Furthermore, our fine-grained noise refinement proves effective, boosting performance by up to 5%. Code is available at \href{https://github.com/hadi-hosseini/noise-refinement}{https://github.com/hadi-hosseini/noise-refinement}.
title Fine-Grained Alignment and Noise Refinement for Compositional Text-to-Image Generation
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2503.06506