Self-Rewarding Large Vision-Language Models for Optimizing Prompts in Text-to-Image Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yang, Hongji, Zhou, Yucheng, Han, Wencheng, Shen, Jianbing
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914201035866112
author Yang, Hongji
Zhou, Yucheng
Han, Wencheng
Shen, Jianbing
author_facet Yang, Hongji
Zhou, Yucheng
Han, Wencheng
Shen, Jianbing
contents Text-to-image models are powerful for producing high-quality images based on given text prompts, but crafting these prompts often requires specialized vocabulary. To address this, existing methods train rewriting models with supervision from large amounts of manually annotated data and trained aesthetic assessment models. To alleviate the dependence on data scale for model training and the biases introduced by trained models, we propose a novel prompt optimization framework, designed to rephrase a simple user prompt into a sophisticated prompt to a text-to-image model. Specifically, we employ the large vision language models (LVLMs) as the solver to rewrite the user prompt, and concurrently, employ LVLMs as a reward model to score the aesthetics and alignment of the images generated by the optimized prompt. Instead of laborious human feedback, we exploit the prior knowledge of the LVLM to provide rewards, i.e., AI feedback. Simultaneously, the solver and the reward model are unified into one model and iterated in reinforcement learning to achieve self-improvement by giving a solution and judging itself. Results on two popular datasets demonstrate that our method outperforms other strong competitors.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16763
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Self-Rewarding Large Vision-Language Models for Optimizing Prompts in Text-to-Image Generation
Yang, Hongji
Zhou, Yucheng
Han, Wencheng
Shen, Jianbing
Computer Vision and Pattern Recognition
Text-to-image models are powerful for producing high-quality images based on given text prompts, but crafting these prompts often requires specialized vocabulary. To address this, existing methods train rewriting models with supervision from large amounts of manually annotated data and trained aesthetic assessment models. To alleviate the dependence on data scale for model training and the biases introduced by trained models, we propose a novel prompt optimization framework, designed to rephrase a simple user prompt into a sophisticated prompt to a text-to-image model. Specifically, we employ the large vision language models (LVLMs) as the solver to rewrite the user prompt, and concurrently, employ LVLMs as a reward model to score the aesthetics and alignment of the images generated by the optimized prompt. Instead of laborious human feedback, we exploit the prior knowledge of the LVLM to provide rewards, i.e., AI feedback. Simultaneously, the solver and the reward model are unified into one model and iterated in reinforcement learning to achieve self-improvement by giving a solution and judging itself. Results on two popular datasets demonstrate that our method outperforms other strong competitors.
title Self-Rewarding Large Vision-Language Models for Optimizing Prompts in Text-to-Image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.16763