Personalized Image Generation with Large Multimodal Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xu, Yiyan, Wang, Wenjie, Zhang, Yang, Tang, Biao, Yan, Peng, Feng, Fuli, He, Xiangnan
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912218410385408
author Xu, Yiyan
Wang, Wenjie
Zhang, Yang
Tang, Biao
Yan, Peng
Feng, Fuli
He, Xiangnan
author_facet Xu, Yiyan
Wang, Wenjie
Zhang, Yang
Tang, Biao
Yan, Peng
Feng, Fuli
He, Xiangnan
contents Personalized content filtering, such as recommender systems, has become a critical infrastructure to alleviate information overload. However, these systems merely filter existing content and are constrained by its limited diversity, making it difficult to meet users' varied content needs. To address this limitation, personalized content generation has emerged as a promising direction with broad applications. Nevertheless, most existing research focuses on personalized text generation, with relatively little attention given to personalized image generation. The limited work in personalized image generation faces challenges in accurately capturing users' visual preferences and needs from noisy user-interacted images and complex multimodal instructions. Worse still, there is a lack of supervised data for training personalized image generation models. To overcome the challenges, we propose a Personalized Image Generation Framework named Pigeon, which adopts exceptional large multimodal models with three dedicated modules to capture users' visual preferences and needs from noisy user history and multimodal instructions. To alleviate the data scarcity, we introduce a two-stage preference alignment scheme, comprising masked preference reconstruction and pairwise preference alignment, to align Pigeon with the personalized image generation task. We apply Pigeon to personalized sticker and movie poster generation, where extensive quantitative results and human evaluation highlight its superiority over various generative baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2410_14170
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Personalized Image Generation with Large Multimodal Models
Xu, Yiyan
Wang, Wenjie
Zhang, Yang
Tang, Biao
Yan, Peng
Feng, Fuli
He, Xiangnan
Information Retrieval
Artificial Intelligence
Multimedia
Personalized content filtering, such as recommender systems, has become a critical infrastructure to alleviate information overload. However, these systems merely filter existing content and are constrained by its limited diversity, making it difficult to meet users' varied content needs. To address this limitation, personalized content generation has emerged as a promising direction with broad applications. Nevertheless, most existing research focuses on personalized text generation, with relatively little attention given to personalized image generation. The limited work in personalized image generation faces challenges in accurately capturing users' visual preferences and needs from noisy user-interacted images and complex multimodal instructions. Worse still, there is a lack of supervised data for training personalized image generation models. To overcome the challenges, we propose a Personalized Image Generation Framework named Pigeon, which adopts exceptional large multimodal models with three dedicated modules to capture users' visual preferences and needs from noisy user history and multimodal instructions. To alleviate the data scarcity, we introduce a two-stage preference alignment scheme, comprising masked preference reconstruction and pairwise preference alignment, to align Pigeon with the personalized image generation task. We apply Pigeon to personalized sticker and movie poster generation, where extensive quantitative results and human evaluation highlight its superiority over various generative baselines.
title Personalized Image Generation with Large Multimodal Models
topic Information Retrieval
Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2410.14170