InstructBooth: Instruction-following Personalized Text-to-Image Generation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chae, Daewon, Park, Nokyung, Kim, Jinkyu, Lee, Kimin
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913234797199360
author Chae, Daewon
Park, Nokyung
Kim, Jinkyu
Lee, Kimin
author_facet Chae, Daewon
Park, Nokyung
Kim, Jinkyu
Lee, Kimin
contents Personalizing text-to-image models using a limited set of images for a specific object has been explored in subject-specific image generation. However, existing methods often face challenges in aligning with text prompts due to overfitting to the limited training images. In this work, we introduce InstructBooth, a novel method designed to enhance image-text alignment in personalized text-to-image models without sacrificing the personalization ability. Our approach first personalizes text-to-image models with a small number of subject-specific images using a unique identifier. After personalization, we fine-tune personalized text-to-image models using reinforcement learning to maximize a reward that quantifies image-text alignment. Additionally, we propose complementary techniques to increase the synergy between these two processes. Our method demonstrates superior image-text alignment compared to existing baselines, while maintaining high personalization ability. In human evaluations, InstructBooth outperforms them when considering all comprehensive factors. Our project page is at https://sites.google.com/view/instructbooth.
format Preprint
id arxiv_https___arxiv_org_abs_2312_03011
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle InstructBooth: Instruction-following Personalized Text-to-Image Generation
Chae, Daewon
Park, Nokyung
Kim, Jinkyu
Lee, Kimin
Computer Vision and Pattern Recognition
Artificial Intelligence
Personalizing text-to-image models using a limited set of images for a specific object has been explored in subject-specific image generation. However, existing methods often face challenges in aligning with text prompts due to overfitting to the limited training images. In this work, we introduce InstructBooth, a novel method designed to enhance image-text alignment in personalized text-to-image models without sacrificing the personalization ability. Our approach first personalizes text-to-image models with a small number of subject-specific images using a unique identifier. After personalization, we fine-tune personalized text-to-image models using reinforcement learning to maximize a reward that quantifies image-text alignment. Additionally, we propose complementary techniques to increase the synergy between these two processes. Our method demonstrates superior image-text alignment compared to existing baselines, while maintaining high personalization ability. In human evaluations, InstructBooth outperforms them when considering all comprehensive factors. Our project page is at https://sites.google.com/view/instructbooth.
title InstructBooth: Instruction-following Personalized Text-to-Image Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2312.03011