Unlocking the Potential of Text-to-Image Diffusion with PAC-Bayesian Theory

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Jiang, Eric Hanchen, Zhang, Yasi, Zhang, Zhi, Wan, Yixin, Lizarraga, Andrew, Li, Shufan, Wu, Ying Nian
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929605868257280
author Jiang, Eric Hanchen
Zhang, Yasi
Zhang, Zhi
Wan, Yixin
Lizarraga, Andrew
Li, Shufan
Wu, Ying Nian
author_facet Jiang, Eric Hanchen
Zhang, Yasi
Zhang, Zhi
Wan, Yixin
Lizarraga, Andrew
Li, Shufan
Wu, Ying Nian
contents Text-to-image (T2I) diffusion models have revolutionized generative modeling by producing high-fidelity, diverse, and visually realistic images from textual prompts. Despite these advances, existing models struggle with complex prompts involving multiple objects and attributes, often misaligning modifiers with their corresponding nouns or neglecting certain elements. Recent attention-based methods have improved object inclusion and linguistic binding, but still face challenges such as attribute misbinding and a lack of robust generalization guarantees. Leveraging the PAC-Bayes framework, we propose a Bayesian approach that designs custom priors over attention distributions to enforce desirable properties, including divergence between objects, alignment between modifiers and their corresponding nouns, minimal attention to irrelevant tokens, and regularization for better generalization. Our approach treats the attention mechanism as an interpretable component, enabling fine-grained control and improved attribute-object alignment. We demonstrate the effectiveness of our method on standard benchmarks, achieving state-of-the-art results across multiple metrics. By integrating custom priors into the denoising process, our method enhances image quality and addresses long-standing challenges in T2I diffusion models, paving the way for more reliable and interpretable generative models.
format Preprint
id arxiv_https___arxiv_org_abs_2411_17472
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Unlocking the Potential of Text-to-Image Diffusion with PAC-Bayesian Theory
Jiang, Eric Hanchen
Zhang, Yasi
Zhang, Zhi
Wan, Yixin
Lizarraga, Andrew
Li, Shufan
Wu, Ying Nian
Computer Vision and Pattern Recognition
Machine Learning
Text-to-image (T2I) diffusion models have revolutionized generative modeling by producing high-fidelity, diverse, and visually realistic images from textual prompts. Despite these advances, existing models struggle with complex prompts involving multiple objects and attributes, often misaligning modifiers with their corresponding nouns or neglecting certain elements. Recent attention-based methods have improved object inclusion and linguistic binding, but still face challenges such as attribute misbinding and a lack of robust generalization guarantees. Leveraging the PAC-Bayes framework, we propose a Bayesian approach that designs custom priors over attention distributions to enforce desirable properties, including divergence between objects, alignment between modifiers and their corresponding nouns, minimal attention to irrelevant tokens, and regularization for better generalization. Our approach treats the attention mechanism as an interpretable component, enabling fine-grained control and improved attribute-object alignment. We demonstrate the effectiveness of our method on standard benchmarks, achieving state-of-the-art results across multiple metrics. By integrating custom priors into the denoising process, our method enhances image quality and addresses long-standing challenges in T2I diffusion models, paving the way for more reliable and interpretable generative models.
title Unlocking the Potential of Text-to-Image Diffusion with PAC-Bayesian Theory
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2411.17472