Learning from Mistakes: Iterative Prompt Relabeling for Text-to-Image Diffusion Model Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Xinyan, Ge, Jiaxin, Zhang, Tianjun, Liu, Jiaming, Zhang, Shanghang
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908273782816768
author Chen, Xinyan
Ge, Jiaxin
Zhang, Tianjun
Liu, Jiaming
Zhang, Shanghang
author_facet Chen, Xinyan
Ge, Jiaxin
Zhang, Tianjun
Liu, Jiaming
Zhang, Shanghang
contents Diffusion models have shown impressive performance in many domains. However, the model's capability to follow natural language instructions (e.g., spatial relationships between objects, generating complex scenes) is still unsatisfactory. In this work, we propose Iterative Prompt Relabeling (IPR), a novel algorithm that aligns images to text through iterative image sampling and prompt relabeling with feedback. IPR first samples a batch of images conditioned on the text, then relabels the text prompts of unmatched text-image pairs with classifier feedback. We conduct thorough experiments on SDv2 and SDXL, testing their capability to follow instructions on spatial relations. With IPR, we improved up to 15.22% (absolute improvement) on the challenging spatial relation VISOR benchmark, demonstrating superior performance compared to previous RL methods. Our code is publicly available at https://github.com/xinyan-cxy/IPR-RLDF.
format Preprint
id arxiv_https___arxiv_org_abs_2312_16204
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Learning from Mistakes: Iterative Prompt Relabeling for Text-to-Image Diffusion Model Training
Chen, Xinyan
Ge, Jiaxin
Zhang, Tianjun
Liu, Jiaming
Zhang, Shanghang
Computer Vision and Pattern Recognition
Diffusion models have shown impressive performance in many domains. However, the model's capability to follow natural language instructions (e.g., spatial relationships between objects, generating complex scenes) is still unsatisfactory. In this work, we propose Iterative Prompt Relabeling (IPR), a novel algorithm that aligns images to text through iterative image sampling and prompt relabeling with feedback. IPR first samples a batch of images conditioned on the text, then relabels the text prompts of unmatched text-image pairs with classifier feedback. We conduct thorough experiments on SDv2 and SDXL, testing their capability to follow instructions on spatial relations. With IPR, we improved up to 15.22% (absolute improvement) on the challenging spatial relation VISOR benchmark, demonstrating superior performance compared to previous RL methods. Our code is publicly available at https://github.com/xinyan-cxy/IPR-RLDF.
title Learning from Mistakes: Iterative Prompt Relabeling for Text-to-Image Diffusion Model Training
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.16204