ArtAug: Enhancing Text-to-Image Generation through Synthesis-Understanding Interaction

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Duan, Zhongjie, Zhao, Qianyi, Chen, Cen, Chen, Daoyuan, Zhou, Wenmeng, Li, Yaliang, Chen, Yingda
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929637380063232
author Duan, Zhongjie
Zhao, Qianyi
Chen, Cen
Chen, Daoyuan
Zhou, Wenmeng
Li, Yaliang
Chen, Yingda
author_facet Duan, Zhongjie
Zhao, Qianyi
Chen, Cen
Chen, Daoyuan
Zhou, Wenmeng
Li, Yaliang
Chen, Yingda
contents The emergence of diffusion models has significantly advanced image synthesis. The recent studies of model interaction and self-corrective reasoning approach in large language models offer new insights for enhancing text-to-image models. Inspired by these studies, we propose a novel method called ArtAug for enhancing text-to-image models in this paper. To the best of our knowledge, ArtAug is the first one that improves image synthesis models via model interactions with understanding models. In the interactions, we leverage human preferences implicitly learned by image understanding models to provide fine-grained suggestions for image synthesis models. The interactions can modify the image content to make it aesthetically pleasing, such as adjusting exposure, changing shooting angles, and adding atmospheric effects. The enhancements brought by the interaction are iteratively fused into the synthesis model itself through an additional enhancement module. This enables the synthesis model to directly produce aesthetically pleasing images without any extra computational cost. In the experiments, we train the ArtAug enhancement module on existing text-to-image models. Various evaluation metrics consistently demonstrate that ArtAug enhances the generative capabilities of text-to-image models without incurring additional computational costs. The source code and models will be released publicly.
format Preprint
id arxiv_https___arxiv_org_abs_2412_12888
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ArtAug: Enhancing Text-to-Image Generation through Synthesis-Understanding Interaction
Duan, Zhongjie
Zhao, Qianyi
Chen, Cen
Chen, Daoyuan
Zhou, Wenmeng
Li, Yaliang
Chen, Yingda
Computer Vision and Pattern Recognition
Artificial Intelligence
The emergence of diffusion models has significantly advanced image synthesis. The recent studies of model interaction and self-corrective reasoning approach in large language models offer new insights for enhancing text-to-image models. Inspired by these studies, we propose a novel method called ArtAug for enhancing text-to-image models in this paper. To the best of our knowledge, ArtAug is the first one that improves image synthesis models via model interactions with understanding models. In the interactions, we leverage human preferences implicitly learned by image understanding models to provide fine-grained suggestions for image synthesis models. The interactions can modify the image content to make it aesthetically pleasing, such as adjusting exposure, changing shooting angles, and adding atmospheric effects. The enhancements brought by the interaction are iteratively fused into the synthesis model itself through an additional enhancement module. This enables the synthesis model to directly produce aesthetically pleasing images without any extra computational cost. In the experiments, we train the ArtAug enhancement module on existing text-to-image models. Various evaluation metrics consistently demonstrate that ArtAug enhances the generative capabilities of text-to-image models without incurring additional computational costs. The source code and models will be released publicly.
title ArtAug: Enhancing Text-to-Image Generation through Synthesis-Understanding Interaction
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2412.12888