VersaT2I: Improving Text-to-Image Models with Versatile Reward

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Jianshu, Chai, Wenhao, Deng, Jie, Huang, Hsiang-Wei, Ye, Tian, Xu, Yichen, Zhang, Jiawei, Hwang, Jenq-Neng, Wang, Gaoang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914730395828224
author Guo, Jianshu
Chai, Wenhao
Deng, Jie
Huang, Hsiang-Wei
Ye, Tian
Xu, Yichen
Zhang, Jiawei
Hwang, Jenq-Neng
Wang, Gaoang
author_facet Guo, Jianshu
Chai, Wenhao
Deng, Jie
Huang, Hsiang-Wei
Ye, Tian
Xu, Yichen
Zhang, Jiawei
Hwang, Jenq-Neng
Wang, Gaoang
contents Recent text-to-image (T2I) models have benefited from large-scale and high-quality data, demonstrating impressive performance. However, these T2I models still struggle to produce images that are aesthetically pleasing, geometrically accurate, faithful to text, and of good low-level quality. We present VersaT2I, a versatile training framework that can boost the performance with multiple rewards of any T2I model. We decompose the quality of the image into several aspects such as aesthetics, text-image alignment, geometry, low-level quality, etc. Then, for every quality aspect, we select high-quality images in this aspect generated by the model as the training set to finetune the T2I model using the Low-Rank Adaptation (LoRA). Furthermore, we introduce a gating function to combine multiple quality aspects, which can avoid conflicts between different quality aspects. Our method is easy to extend and does not require any manual annotation, reinforcement learning, or model architecture changes. Extensive experiments demonstrate that VersaT2I outperforms the baseline methods across various quality criteria.
format Preprint
id arxiv_https___arxiv_org_abs_2403_18493
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VersaT2I: Improving Text-to-Image Models with Versatile Reward
Guo, Jianshu
Chai, Wenhao
Deng, Jie
Huang, Hsiang-Wei
Ye, Tian
Xu, Yichen
Zhang, Jiawei
Hwang, Jenq-Neng
Wang, Gaoang
Computer Vision and Pattern Recognition
Recent text-to-image (T2I) models have benefited from large-scale and high-quality data, demonstrating impressive performance. However, these T2I models still struggle to produce images that are aesthetically pleasing, geometrically accurate, faithful to text, and of good low-level quality. We present VersaT2I, a versatile training framework that can boost the performance with multiple rewards of any T2I model. We decompose the quality of the image into several aspects such as aesthetics, text-image alignment, geometry, low-level quality, etc. Then, for every quality aspect, we select high-quality images in this aspect generated by the model as the training set to finetune the T2I model using the Low-Rank Adaptation (LoRA). Furthermore, we introduce a gating function to combine multiple quality aspects, which can avoid conflicts between different quality aspects. Our method is easy to extend and does not require any manual annotation, reinforcement learning, or model architecture changes. Extensive experiments demonstrate that VersaT2I outperforms the baseline methods across various quality criteria.
title VersaT2I: Improving Text-to-Image Models with Versatile Reward
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.18493