POET: Supporting Prompting Creativity and Personalization with Automated Expansion of Text-to-Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Han, Evans Xu, Zhang, Alice Qian, Zhu, Haiyi, Shen, Hong, Liang, Paul Pu, Hsieh, Jane
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916932783964160
author Han, Evans Xu
Zhang, Alice Qian
Zhu, Haiyi
Shen, Hong
Liang, Paul Pu
Hsieh, Jane
author_facet Han, Evans Xu
Zhang, Alice Qian
Zhu, Haiyi
Shen, Hong
Liang, Paul Pu
Hsieh, Jane
contents State-of-the-art visual generative AI tools hold immense potential to assist users in the early ideation stages of creative tasks -- offering the ability to generate (rather than search for) novel and unprecedented (instead of existing) images of considerable quality that also adhere to boundless combinations of user specifications. However, many large-scale text-to-image systems are designed for broad applicability, yielding conventional output that may limit creative exploration. They also employ interaction methods that may be difficult for beginners. Given that creative end users often operate in diverse, context-specific ways that are often unpredictable, more variation and personalization are necessary. We introduce POET, a real-time interactive tool that (1) automatically discovers dimensions of homogeneity in text-to-image generative models, (2) expands these dimensions to diversify the output space of generated images, and (3) learns from user feedback to personalize expansions. An evaluation with 28 users spanning four creative task domains demonstrated POET's ability to generate results with higher perceived diversity and help users reach satisfaction in fewer prompts during creative tasks, thereby prompting them to deliberate and reflect more on a wider range of possible produced results during the co-creative process. Focusing on visual creativity, POET offers a first glimpse of how interaction techniques of future text-to-image generation tools may support and align with more pluralistic values and the needs of end users during the ideation stages of their work.
format Preprint
id arxiv_https___arxiv_org_abs_2504_13392
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle POET: Supporting Prompting Creativity and Personalization with Automated Expansion of Text-to-Image Generation
Han, Evans Xu
Zhang, Alice Qian
Zhu, Haiyi
Shen, Hong
Liang, Paul Pu
Hsieh, Jane
Computer Vision and Pattern Recognition
Human-Computer Interaction
State-of-the-art visual generative AI tools hold immense potential to assist users in the early ideation stages of creative tasks -- offering the ability to generate (rather than search for) novel and unprecedented (instead of existing) images of considerable quality that also adhere to boundless combinations of user specifications. However, many large-scale text-to-image systems are designed for broad applicability, yielding conventional output that may limit creative exploration. They also employ interaction methods that may be difficult for beginners. Given that creative end users often operate in diverse, context-specific ways that are often unpredictable, more variation and personalization are necessary. We introduce POET, a real-time interactive tool that (1) automatically discovers dimensions of homogeneity in text-to-image generative models, (2) expands these dimensions to diversify the output space of generated images, and (3) learns from user feedback to personalize expansions. An evaluation with 28 users spanning four creative task domains demonstrated POET's ability to generate results with higher perceived diversity and help users reach satisfaction in fewer prompts during creative tasks, thereby prompting them to deliberate and reflect more on a wider range of possible produced results during the co-creative process. Focusing on visual creativity, POET offers a first glimpse of how interaction techniques of future text-to-image generation tools may support and align with more pluralistic values and the needs of end users during the ideation stages of their work.
title POET: Supporting Prompting Creativity and Personalization with Automated Expansion of Text-to-Image Generation
topic Computer Vision and Pattern Recognition
Human-Computer Interaction
url https://arxiv.org/abs/2504.13392