Personalized Preference Fine-tuning of Diffusion Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Dang, Meihua, Singh, Anikait, Zhou, Linqi, Ermon, Stefano, Song, Jiaming
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915099198881792
author Dang, Meihua
Singh, Anikait
Zhou, Linqi
Ermon, Stefano
Song, Jiaming
author_facet Dang, Meihua
Singh, Anikait
Zhou, Linqi
Ermon, Stefano
Song, Jiaming
contents RLHF techniques like DPO can significantly improve the generation quality of text-to-image diffusion models. However, these methods optimize for a single reward that aligns model generation with population-level preferences, neglecting the nuances of individual users' beliefs or values. This lack of personalization limits the efficacy of these models. To bridge this gap, we introduce PPD, a multi-reward optimization objective that aligns diffusion models with personalized preferences. With PPD, a diffusion model learns the individual preferences of a population of users in a few-shot way, enabling generalization to unseen users. Specifically, our approach (1) leverages a vision-language model (VLM) to extract personal preference embeddings from a small set of pairwise preference examples, and then (2) incorporates the embeddings into diffusion models through cross attention. Conditioning on user embeddings, the text-to-image models are fine-tuned with the DPO objective, simultaneously optimizing for alignment with the preferences of multiple users. Empirical results demonstrate that our method effectively optimizes for multiple reward functions and can interpolate between them during inference. In real-world user scenarios, with as few as four preference examples from a new user, our approach achieves an average win rate of 76\% over Stable Cascade, generating images that more accurately reflect specific user preferences.
format Preprint
id arxiv_https___arxiv_org_abs_2501_06655
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Personalized Preference Fine-tuning of Diffusion Models
Dang, Meihua
Singh, Anikait
Zhou, Linqi
Ermon, Stefano
Song, Jiaming
Machine Learning
Computer Vision and Pattern Recognition
RLHF techniques like DPO can significantly improve the generation quality of text-to-image diffusion models. However, these methods optimize for a single reward that aligns model generation with population-level preferences, neglecting the nuances of individual users' beliefs or values. This lack of personalization limits the efficacy of these models. To bridge this gap, we introduce PPD, a multi-reward optimization objective that aligns diffusion models with personalized preferences. With PPD, a diffusion model learns the individual preferences of a population of users in a few-shot way, enabling generalization to unseen users. Specifically, our approach (1) leverages a vision-language model (VLM) to extract personal preference embeddings from a small set of pairwise preference examples, and then (2) incorporates the embeddings into diffusion models through cross attention. Conditioning on user embeddings, the text-to-image models are fine-tuned with the DPO objective, simultaneously optimizing for alignment with the preferences of multiple users. Empirical results demonstrate that our method effectively optimizes for multiple reward functions and can interpolate between them during inference. In real-world user scenarios, with as few as four preference examples from a new user, our approach achieves an average win rate of 76\% over Stable Cascade, generating images that more accurately reflect specific user preferences.
title Personalized Preference Fine-tuning of Diffusion Models
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.06655