Enhancing Intent Understanding for Ambiguous prompt: A Human-Machine Co-Adaption Strategy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Yangfan, Wang, Jianhui, Wang, Yijin, Zhong, Yan, Song, Xinyuan, Lin, Junjiang, Yuan, Xinhang, Tang, Jingqun, Xin, Yi, Zhang, Hao, Li, Yuchen, Zhang, Zijian, He, Hongyang, Xu, Tianxiang, Zhang, Miao, Lu, Kuan, Huo, Menghao, Li, Keqin, Chen, Jiaqi, Shi, Tianyu, Ni, Jianyuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916953796378624
author He, Yangfan
Wang, Jianhui
Wang, Yijin
Zhong, Yan
Song, Xinyuan
Lin, Junjiang
Yuan, Xinhang
Tang, Jingqun
Xin, Yi
Zhang, Hao
Li, Yuchen
Zhang, Zijian
He, Hongyang
Xu, Tianxiang
Zhang, Miao
Lu, Kuan
Huo, Menghao
Li, Keqin
Chen, Jiaqi
Shi, Tianyu
Ni, Jianyuan
author_facet He, Yangfan
Wang, Jianhui
Wang, Yijin
Zhong, Yan
Song, Xinyuan
Lin, Junjiang
Yuan, Xinhang
Tang, Jingqun
Xin, Yi
Zhang, Hao
Li, Yuchen
Zhang, Zijian
He, Hongyang
Xu, Tianxiang
Zhang, Miao
Lu, Kuan
Huo, Menghao
Li, Keqin
Chen, Jiaqi
Shi, Tianyu
Ni, Jianyuan
contents Current image generation systems produce high-quality images but struggle with ambiguous user prompts, making interpretation of actual user intentions difficult. Many users must modify their prompts several times to ensure the generated images meet their expectations. While some methods focus on enhancing prompts to make the generated images fit user needs, the model is still hard to understand users' real needs, especially for non-expert users. In this research, we aim to enhance the visual parameter-tuning process, making the model user-friendly for individuals without specialized knowledge and better understand user needs. We propose a human-machine co-adaption strategy using mutual information between the user's prompts and the pictures under modification as the optimizing target to make the system better adapt to user needs. We find that an improved model can reduce the necessity for multiple rounds of adjustments. We also collect multi-round dialogue datasets with prompts and images pairs and user intent. Various experiments demonstrate the effectiveness of the proposed method in our proposed dataset. Our dataset and annotation tools will be available.
format Preprint
id arxiv_https___arxiv_org_abs_2501_15167
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing Intent Understanding for Ambiguous prompt: A Human-Machine Co-Adaption Strategy
He, Yangfan
Wang, Jianhui
Wang, Yijin
Zhong, Yan
Song, Xinyuan
Lin, Junjiang
Yuan, Xinhang
Tang, Jingqun
Xin, Yi
Zhang, Hao
Li, Yuchen
Zhang, Zijian
He, Hongyang
Xu, Tianxiang
Zhang, Miao
Lu, Kuan
Huo, Menghao
Li, Keqin
Chen, Jiaqi
Shi, Tianyu
Ni, Jianyuan
Computer Vision and Pattern Recognition
Current image generation systems produce high-quality images but struggle with ambiguous user prompts, making interpretation of actual user intentions difficult. Many users must modify their prompts several times to ensure the generated images meet their expectations. While some methods focus on enhancing prompts to make the generated images fit user needs, the model is still hard to understand users' real needs, especially for non-expert users. In this research, we aim to enhance the visual parameter-tuning process, making the model user-friendly for individuals without specialized knowledge and better understand user needs. We propose a human-machine co-adaption strategy using mutual information between the user's prompts and the pictures under modification as the optimizing target to make the system better adapt to user needs. We find that an improved model can reduce the necessity for multiple rounds of adjustments. We also collect multi-round dialogue datasets with prompts and images pairs and user intent. Various experiments demonstrate the effectiveness of the proposed method in our proposed dataset. Our dataset and annotation tools will be available.
title Enhancing Intent Understanding for Ambiguous prompt: A Human-Machine Co-Adaption Strategy
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.15167