Adaptive Prompt Elicitation for Text-to-Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wen, Xinyi, Hegemann, Lena, Jin, Xiaofu, Ma, Shuai, Oulasvirta, Antti
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914494038409216
author Wen, Xinyi
Hegemann, Lena
Jin, Xiaofu
Ma, Shuai
Oulasvirta, Antti
author_facet Wen, Xinyi
Hegemann, Lena
Jin, Xiaofu
Ma, Shuai
Oulasvirta, Antti
contents Aligning text-to-image generation with user intent remains challenging, as users frequently provide ambiguous inputs and struggle with model idiosyncrasies. We propose Adaptive Prompt Elicitation (APE), a technique that adaptively poses visual queries to help users refine prompts without extensive writing. Our technical contribution is a formulation of interactive intent inference under an information-theoretic framework. APE represents latent user intent as interpretable feature requirements using language model priors, adaptively generates visual queries, and compiles elicited requirements into effective prompts. Evaluation on IDEA-Bench and DesignBench shows that APE achieves stronger alignment with improved efficiency. A user study with 128 participants on user-defined tasks demonstrates 19.8% higher perceived alignment without increased workload. Our work contributes a principled approach to prompting that offers an effective and efficient complement to the prevailing prompt-based interaction paradigm with text-to-image models.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04713
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Adaptive Prompt Elicitation for Text-to-Image Generation
Wen, Xinyi
Hegemann, Lena
Jin, Xiaofu
Ma, Shuai
Oulasvirta, Antti
Human-Computer Interaction
Artificial Intelligence
Computer Vision and Pattern Recognition
I.2; I.6
Aligning text-to-image generation with user intent remains challenging, as users frequently provide ambiguous inputs and struggle with model idiosyncrasies. We propose Adaptive Prompt Elicitation (APE), a technique that adaptively poses visual queries to help users refine prompts without extensive writing. Our technical contribution is a formulation of interactive intent inference under an information-theoretic framework. APE represents latent user intent as interpretable feature requirements using language model priors, adaptively generates visual queries, and compiles elicited requirements into effective prompts. Evaluation on IDEA-Bench and DesignBench shows that APE achieves stronger alignment with improved efficiency. A user study with 128 participants on user-defined tasks demonstrates 19.8% higher perceived alignment without increased workload. Our work contributes a principled approach to prompting that offers an effective and efficient complement to the prevailing prompt-based interaction paradigm with text-to-image models.
title Adaptive Prompt Elicitation for Text-to-Image Generation
topic Human-Computer Interaction
Artificial Intelligence
Computer Vision and Pattern Recognition
I.2; I.6
url https://arxiv.org/abs/2602.04713