PromptSculptor: Multi-Agent Based Text-to-Image Prompt Optimization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xiang, Dawei, Xu, Wenyan, Chu, Kexin, Ding, Tianqi, Shen, Zixu, Zeng, Yiming, Su, Jianchang, Zhang, Wei
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914053422579712
author Xiang, Dawei
Xu, Wenyan
Chu, Kexin
Ding, Tianqi
Shen, Zixu
Zeng, Yiming
Su, Jianchang
Zhang, Wei
author_facet Xiang, Dawei
Xu, Wenyan
Chu, Kexin
Ding, Tianqi
Shen, Zixu
Zeng, Yiming
Su, Jianchang
Zhang, Wei
contents The rapid advancement of generative AI has democratized access to powerful tools such as Text-to-Image models. However, to generate high-quality images, users must still craft detailed prompts specifying scene, style, and context-often through multiple rounds of refinement. We propose PromptSculptor, a novel multi-agent framework that automates this iterative prompt optimization process. Our system decomposes the task into four specialized agents that work collaboratively to transform a short, vague user prompt into a comprehensive, refined prompt. By leveraging Chain-of-Thought reasoning, our framework effectively infers hidden context and enriches scene and background details. To iteratively refine the prompt, a self-evaluation agent aligns the modified prompt with the original input, while a feedback-tuning agent incorporates user feedback for further refinement. Experimental results demonstrate that PromptSculptor significantly enhances output quality and reduces the number of iterations needed for user satisfaction. Moreover, its model-agnostic design allows seamless integration with various T2I models, paving the way for industrial applications.
format Preprint
id arxiv_https___arxiv_org_abs_2509_12446
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PromptSculptor: Multi-Agent Based Text-to-Image Prompt Optimization
Xiang, Dawei
Xu, Wenyan
Chu, Kexin
Ding, Tianqi
Shen, Zixu
Zeng, Yiming
Su, Jianchang
Zhang, Wei
Multiagent Systems
Artificial Intelligence
The rapid advancement of generative AI has democratized access to powerful tools such as Text-to-Image models. However, to generate high-quality images, users must still craft detailed prompts specifying scene, style, and context-often through multiple rounds of refinement. We propose PromptSculptor, a novel multi-agent framework that automates this iterative prompt optimization process. Our system decomposes the task into four specialized agents that work collaboratively to transform a short, vague user prompt into a comprehensive, refined prompt. By leveraging Chain-of-Thought reasoning, our framework effectively infers hidden context and enriches scene and background details. To iteratively refine the prompt, a self-evaluation agent aligns the modified prompt with the original input, while a feedback-tuning agent incorporates user feedback for further refinement. Experimental results demonstrate that PromptSculptor significantly enhances output quality and reduces the number of iterations needed for user satisfaction. Moreover, its model-agnostic design allows seamless integration with various T2I models, paving the way for industrial applications.
title PromptSculptor: Multi-Agent Based Text-to-Image Prompt Optimization
topic Multiagent Systems
Artificial Intelligence
url https://arxiv.org/abs/2509.12446