Learning User Preferences for Image Generation Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mo, Wenyi, Ba, Ying, Zhang, Tianyu, Bai, Yalong, Li, Biye
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915440258711552
author Mo, Wenyi
Ba, Ying
Zhang, Tianyu
Bai, Yalong
Li, Biye
author_facet Mo, Wenyi
Ba, Ying
Zhang, Tianyu
Bai, Yalong
Li, Biye
contents User preference prediction requires a comprehensive and accurate understanding of individual tastes. This includes both surface-level attributes, such as color and style, and deeper content-related aspects, such as themes and composition. However, existing methods typically rely on general human preferences or assume static user profiles, often neglecting individual variability and the dynamic, multifaceted nature of personal taste. To address these limitations, we propose an approach built upon Multimodal Large Language Models, introducing contrastive preference loss and preference tokens to learn personalized user preferences from historical interactions. The contrastive preference loss is designed to effectively distinguish between user ''likes'' and ''dislikes'', while the learnable preference tokens capture shared interest representations among existing users, enabling the model to activate group-specific preferences and enhance consistency across similar users. Extensive experiments demonstrate our model outperforms other methods in preference prediction accuracy, effectively identifying users with similar aesthetic inclinations and providing more precise guidance for generating images that align with individual tastes. The project page is \texttt{https://learn-user-pref.github.io/}.
format Preprint
id arxiv_https___arxiv_org_abs_2508_08220
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning User Preferences for Image Generation Model
Mo, Wenyi
Ba, Ying
Zhang, Tianyu
Bai, Yalong
Li, Biye
Computer Vision and Pattern Recognition
User preference prediction requires a comprehensive and accurate understanding of individual tastes. This includes both surface-level attributes, such as color and style, and deeper content-related aspects, such as themes and composition. However, existing methods typically rely on general human preferences or assume static user profiles, often neglecting individual variability and the dynamic, multifaceted nature of personal taste. To address these limitations, we propose an approach built upon Multimodal Large Language Models, introducing contrastive preference loss and preference tokens to learn personalized user preferences from historical interactions. The contrastive preference loss is designed to effectively distinguish between user ''likes'' and ''dislikes'', while the learnable preference tokens capture shared interest representations among existing users, enabling the model to activate group-specific preferences and enhance consistency across similar users. Extensive experiments demonstrate our model outperforms other methods in preference prediction accuracy, effectively identifying users with similar aesthetic inclinations and providing more precise guidance for generating images that align with individual tastes. The project page is \texttt{https://learn-user-pref.github.io/}.
title Learning User Preferences for Image Generation Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.08220