TextTIGER: Text-based Intelligent Generation with Entity Prompt Refinement for Text-to-Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ozaki, Shintaro, Jinno, Tomoyuki, Hayashi, Kazuki, Sakai, Yusuke, Kwon, Jingun, Kamigaito, Hidetaka, Hayashi, Katsuhiko, Okumura, Manabu, Watanabe, Taro
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910140743024640
author Ozaki, Shintaro
Jinno, Tomoyuki
Hayashi, Kazuki
Sakai, Yusuke
Kwon, Jingun
Kamigaito, Hidetaka
Hayashi, Katsuhiko
Okumura, Manabu
Watanabe, Taro
author_facet Ozaki, Shintaro
Jinno, Tomoyuki
Hayashi, Kazuki
Sakai, Yusuke
Kwon, Jingun
Kamigaito, Hidetaka
Hayashi, Katsuhiko
Okumura, Manabu
Watanabe, Taro
contents When generating images from prompts that include specific entities, the model must retain as much entity-specific knowledge as possible. However, the number of entities is almost countless, and new entities emerge; memorizing all of them completely is not realistic. To bridge this gap, our work proposes Text-based Intelligent Generation with Entity Prompt Refinement (TextTIGER). TextTIGER strengthens knowledge about entities that appear in the prompt by augmenting external information and then summarizes the expanded descriptions with large language models, preventing performance degradation that arises from excessively long inputs. To evaluate our method, we construct a new dataset consisting of captions, images, detailed descriptions, and lists of entities. Experiments with multiple image generation models show that TextTIGER improves image generation performance on widely used evaluation metrics compared with prompts that use captions alone. In addition, using Multimodal LLM (MLLM)-as-a-judge, which shows a strong correlation with human evaluation, we demonstrate that our method consistently achieves higher scores, which underscores its effectiveness. These results show that strengthening entity-related descriptions, summarizing them, and refining prompts to an appropriate length leads to substantial improvements in image generation performance. We will release the created dataset and code upon acceptance.
format Preprint
id arxiv_https___arxiv_org_abs_2504_18269
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TextTIGER: Text-based Intelligent Generation with Entity Prompt Refinement for Text-to-Image Generation
Ozaki, Shintaro
Jinno, Tomoyuki
Hayashi, Kazuki
Sakai, Yusuke
Kwon, Jingun
Kamigaito, Hidetaka
Hayashi, Katsuhiko
Okumura, Manabu
Watanabe, Taro
Computation and Language
Computer Vision and Pattern Recognition
When generating images from prompts that include specific entities, the model must retain as much entity-specific knowledge as possible. However, the number of entities is almost countless, and new entities emerge; memorizing all of them completely is not realistic. To bridge this gap, our work proposes Text-based Intelligent Generation with Entity Prompt Refinement (TextTIGER). TextTIGER strengthens knowledge about entities that appear in the prompt by augmenting external information and then summarizes the expanded descriptions with large language models, preventing performance degradation that arises from excessively long inputs. To evaluate our method, we construct a new dataset consisting of captions, images, detailed descriptions, and lists of entities. Experiments with multiple image generation models show that TextTIGER improves image generation performance on widely used evaluation metrics compared with prompts that use captions alone. In addition, using Multimodal LLM (MLLM)-as-a-judge, which shows a strong correlation with human evaluation, we demonstrate that our method consistently achieves higher scores, which underscores its effectiveness. These results show that strengthening entity-related descriptions, summarizing them, and refining prompts to an appropriate length leads to substantial improvements in image generation performance. We will release the created dataset and code upon acceptance.
title TextTIGER: Text-based Intelligent Generation with Entity Prompt Refinement for Text-to-Image Generation
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.18269