TIPO: Text to Image with Text Presampling for Prompt Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yeh, Shih-Ying, Li, Yi, Park, Sang-Hyun, Oh, Giyeong, Wang, Xuehai, Song, Min, Yu, Youngjae, Lai, Shang-Hong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912886165602304
author Yeh, Shih-Ying
Li, Yi
Park, Sang-Hyun
Oh, Giyeong
Wang, Xuehai
Song, Min
Yu, Youngjae
Lai, Shang-Hong
author_facet Yeh, Shih-Ying
Li, Yi
Park, Sang-Hyun
Oh, Giyeong
Wang, Xuehai
Song, Min
Yu, Youngjae
Lai, Shang-Hong
contents TIPO (Text-to-Image Prompt Optimization) introduces an efficient approach for automatic prompt refinement in text-to-image (T2I) generation. Starting from simple user prompts, TIPO leverages a lightweight pre-trained model to expand these prompts into richer and more detailed versions. Conceptually, TIPO samples refined prompts from a targeted sub-distribution within the broader semantic space, preserving the original intent while significantly improving visual quality, coherence, and detail. Unlike resource-intensive methods based on large language models (LLMs) or reinforcement learning (RL), TIPO offers strong computational efficiency and scalability, opening new possibilities for effective automated prompt engineering in T2I tasks. Extensive experiments across multiple domains demonstrate that TIPO achieves stronger text alignment, reduced visual artifacts, and consistently higher human preference rates, while maintaining competitive aesthetic quality. These results highlight the effectiveness of distribution-aligned prompt engineering and point toward broader opportunities for scalable, automated refinement in text-to-image generation.
format Preprint
id arxiv_https___arxiv_org_abs_2411_08127
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TIPO: Text to Image with Text Presampling for Prompt Optimization
Yeh, Shih-Ying
Li, Yi
Park, Sang-Hyun
Oh, Giyeong
Wang, Xuehai
Song, Min
Yu, Youngjae
Lai, Shang-Hong
Computer Vision and Pattern Recognition
TIPO (Text-to-Image Prompt Optimization) introduces an efficient approach for automatic prompt refinement in text-to-image (T2I) generation. Starting from simple user prompts, TIPO leverages a lightweight pre-trained model to expand these prompts into richer and more detailed versions. Conceptually, TIPO samples refined prompts from a targeted sub-distribution within the broader semantic space, preserving the original intent while significantly improving visual quality, coherence, and detail. Unlike resource-intensive methods based on large language models (LLMs) or reinforcement learning (RL), TIPO offers strong computational efficiency and scalability, opening new possibilities for effective automated prompt engineering in T2I tasks. Extensive experiments across multiple domains demonstrate that TIPO achieves stronger text alignment, reduced visual artifacts, and consistently higher human preference rates, while maintaining competitive aesthetic quality. These results highlight the effectiveness of distribution-aligned prompt engineering and point toward broader opportunities for scalable, automated refinement in text-to-image generation.
title TIPO: Text to Image with Text Presampling for Prompt Optimization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.08127