Parrot: Pareto-optimal Multi-Reward Reinforcement Learning Framework for Text-to-Image Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lee, Seung Hyun, Li, Yinxiao, Ke, Junjie, Yoo, Innfarn, Zhang, Han, Yu, Jiahui, Wang, Qifei, Deng, Fei, Entis, Glenn, He, Junfeng, Li, Gang, Kim, Sangpil, Essa, Irfan, Yang, Feng
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916323864346624
author Lee, Seung Hyun
Li, Yinxiao
Ke, Junjie
Yoo, Innfarn
Zhang, Han
Yu, Jiahui
Wang, Qifei
Deng, Fei
Entis, Glenn
He, Junfeng
Li, Gang
Kim, Sangpil
Essa, Irfan
Yang, Feng
author_facet Lee, Seung Hyun
Li, Yinxiao
Ke, Junjie
Yoo, Innfarn
Zhang, Han
Yu, Jiahui
Wang, Qifei
Deng, Fei
Entis, Glenn
He, Junfeng
Li, Gang
Kim, Sangpil
Essa, Irfan
Yang, Feng
contents Recent works have demonstrated that using reinforcement learning (RL) with multiple quality rewards can improve the quality of generated images in text-to-image (T2I) generation. However, manually adjusting reward weights poses challenges and may cause over-optimization in certain metrics. To solve this, we propose Parrot, which addresses the issue through multi-objective optimization and introduces an effective multi-reward optimization strategy to approximate Pareto optimal. Utilizing batch-wise Pareto optimal selection, Parrot automatically identifies the optimal trade-off among different rewards. We use the novel multi-reward optimization algorithm to jointly optimize the T2I model and a prompt expansion network, resulting in significant improvement of image quality and also allow to control the trade-off of different rewards using a reward related prompt during inference. Furthermore, we introduce original prompt-centered guidance at inference time, ensuring fidelity to user input after prompt expansion. Extensive experiments and a user study validate the superiority of Parrot over several baselines across various quality criteria, including aesthetics, human preference, text-image alignment, and image sentiment.
format Preprint
id arxiv_https___arxiv_org_abs_2401_05675
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Parrot: Pareto-optimal Multi-Reward Reinforcement Learning Framework for Text-to-Image Generation
Lee, Seung Hyun
Li, Yinxiao
Ke, Junjie
Yoo, Innfarn
Zhang, Han
Yu, Jiahui
Wang, Qifei
Deng, Fei
Entis, Glenn
He, Junfeng
Li, Gang
Kim, Sangpil
Essa, Irfan
Yang, Feng
Computer Vision and Pattern Recognition
Recent works have demonstrated that using reinforcement learning (RL) with multiple quality rewards can improve the quality of generated images in text-to-image (T2I) generation. However, manually adjusting reward weights poses challenges and may cause over-optimization in certain metrics. To solve this, we propose Parrot, which addresses the issue through multi-objective optimization and introduces an effective multi-reward optimization strategy to approximate Pareto optimal. Utilizing batch-wise Pareto optimal selection, Parrot automatically identifies the optimal trade-off among different rewards. We use the novel multi-reward optimization algorithm to jointly optimize the T2I model and a prompt expansion network, resulting in significant improvement of image quality and also allow to control the trade-off of different rewards using a reward related prompt during inference. Furthermore, we introduce original prompt-centered guidance at inference time, ensuring fidelity to user input after prompt expansion. Extensive experiments and a user study validate the superiority of Parrot over several baselines across various quality criteria, including aesthetics, human preference, text-image alignment, and image sentiment.
title Parrot: Pareto-optimal Multi-Reward Reinforcement Learning Framework for Text-to-Image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2401.05675