Enhancing Intent Understanding for Ambiguous prompt: A Human-Machine Co-Adaption Strategy
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916953796378624 |
|---|---|
| author | He, Yangfan Wang, Jianhui Wang, Yijin Zhong, Yan Song, Xinyuan Lin, Junjiang Yuan, Xinhang Tang, Jingqun Xin, Yi Zhang, Hao Li, Yuchen Zhang, Zijian He, Hongyang Xu, Tianxiang Zhang, Miao Lu, Kuan Huo, Menghao Li, Keqin Chen, Jiaqi Shi, Tianyu Ni, Jianyuan |
| author_facet | He, Yangfan Wang, Jianhui Wang, Yijin Zhong, Yan Song, Xinyuan Lin, Junjiang Yuan, Xinhang Tang, Jingqun Xin, Yi Zhang, Hao Li, Yuchen Zhang, Zijian He, Hongyang Xu, Tianxiang Zhang, Miao Lu, Kuan Huo, Menghao Li, Keqin Chen, Jiaqi Shi, Tianyu Ni, Jianyuan |
| contents | Current image generation systems produce high-quality images but struggle with ambiguous user prompts, making interpretation of actual user intentions difficult. Many users must modify their prompts several times to ensure the generated images meet their expectations. While some methods focus on enhancing prompts to make the generated images fit user needs, the model is still hard to understand users' real needs, especially for non-expert users. In this research, we aim to enhance the visual parameter-tuning process, making the model user-friendly for individuals without specialized knowledge and better understand user needs. We propose a human-machine co-adaption strategy using mutual information between the user's prompts and the pictures under modification as the optimizing target to make the system better adapt to user needs. We find that an improved model can reduce the necessity for multiple rounds of adjustments. We also collect multi-round dialogue datasets with prompts and images pairs and user intent. Various experiments demonstrate the effectiveness of the proposed method in our proposed dataset. Our dataset and annotation tools will be available. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_15167 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Enhancing Intent Understanding for Ambiguous prompt: A Human-Machine Co-Adaption Strategy He, Yangfan Wang, Jianhui Wang, Yijin Zhong, Yan Song, Xinyuan Lin, Junjiang Yuan, Xinhang Tang, Jingqun Xin, Yi Zhang, Hao Li, Yuchen Zhang, Zijian He, Hongyang Xu, Tianxiang Zhang, Miao Lu, Kuan Huo, Menghao Li, Keqin Chen, Jiaqi Shi, Tianyu Ni, Jianyuan Computer Vision and Pattern Recognition Current image generation systems produce high-quality images but struggle with ambiguous user prompts, making interpretation of actual user intentions difficult. Many users must modify their prompts several times to ensure the generated images meet their expectations. While some methods focus on enhancing prompts to make the generated images fit user needs, the model is still hard to understand users' real needs, especially for non-expert users. In this research, we aim to enhance the visual parameter-tuning process, making the model user-friendly for individuals without specialized knowledge and better understand user needs. We propose a human-machine co-adaption strategy using mutual information between the user's prompts and the pictures under modification as the optimizing target to make the system better adapt to user needs. We find that an improved model can reduce the necessity for multiple rounds of adjustments. We also collect multi-round dialogue datasets with prompts and images pairs and user intent. Various experiments demonstrate the effectiveness of the proposed method in our proposed dataset. Our dataset and annotation tools will be available. |
| title | Enhancing Intent Understanding for Ambiguous prompt: A Human-Machine Co-Adaption Strategy |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2501.15167 |