Adaptive Prompt Elicitation for Text-to-Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914494038409216 |
|---|---|
| author | Wen, Xinyi Hegemann, Lena Jin, Xiaofu Ma, Shuai Oulasvirta, Antti |
| author_facet | Wen, Xinyi Hegemann, Lena Jin, Xiaofu Ma, Shuai Oulasvirta, Antti |
| contents | Aligning text-to-image generation with user intent remains challenging, as users frequently provide ambiguous inputs and struggle with model idiosyncrasies. We propose Adaptive Prompt Elicitation (APE), a technique that adaptively poses visual queries to help users refine prompts without extensive writing. Our technical contribution is a formulation of interactive intent inference under an information-theoretic framework. APE represents latent user intent as interpretable feature requirements using language model priors, adaptively generates visual queries, and compiles elicited requirements into effective prompts. Evaluation on IDEA-Bench and DesignBench shows that APE achieves stronger alignment with improved efficiency. A user study with 128 participants on user-defined tasks demonstrates 19.8% higher perceived alignment without increased workload. Our work contributes a principled approach to prompting that offers an effective and efficient complement to the prevailing prompt-based interaction paradigm with text-to-image models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_04713 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Adaptive Prompt Elicitation for Text-to-Image Generation Wen, Xinyi Hegemann, Lena Jin, Xiaofu Ma, Shuai Oulasvirta, Antti Human-Computer Interaction Artificial Intelligence Computer Vision and Pattern Recognition I.2; I.6 Aligning text-to-image generation with user intent remains challenging, as users frequently provide ambiguous inputs and struggle with model idiosyncrasies. We propose Adaptive Prompt Elicitation (APE), a technique that adaptively poses visual queries to help users refine prompts without extensive writing. Our technical contribution is a formulation of interactive intent inference under an information-theoretic framework. APE represents latent user intent as interpretable feature requirements using language model priors, adaptively generates visual queries, and compiles elicited requirements into effective prompts. Evaluation on IDEA-Bench and DesignBench shows that APE achieves stronger alignment with improved efficiency. A user study with 128 participants on user-defined tasks demonstrates 19.8% higher perceived alignment without increased workload. Our work contributes a principled approach to prompting that offers an effective and efficient complement to the prevailing prompt-based interaction paradigm with text-to-image models. |
| title | Adaptive Prompt Elicitation for Text-to-Image Generation |
| topic | Human-Computer Interaction Artificial Intelligence Computer Vision and Pattern Recognition I.2; I.6 |
| url | https://arxiv.org/abs/2602.04713 |