Margin-aware Preference Optimization for Aligning Diffusion Models without Reference

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hong, Jiwoo, Paul, Sayak, Lee, Noah, Rasul, Kashif, Thorne, James, Jeong, Jongheon
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917120349044736
author Hong, Jiwoo
Paul, Sayak
Lee, Noah
Rasul, Kashif
Thorne, James
Jeong, Jongheon
author_facet Hong, Jiwoo
Paul, Sayak
Lee, Noah
Rasul, Kashif
Thorne, James
Jeong, Jongheon
contents Modern preference alignment methods, such as DPO, rely on divergence regularization to a reference model for training stability-but this creates a fundamental problem we call "reference mismatch." In this paper, we investigate the negative impacts of reference mismatch in aligning text-to-image (T2I) diffusion models, showing that larger reference mismatch hinders effective adaptation given the same amount of data, e.g., as when learning new artistic styles, or personalizing to specific objects. We demonstrate this phenomenon across text-to-image (T2I) diffusion models and introduce margin-aware preference optimization (MaPO), a reference-agnostic approach that breaks free from this constraint. By directly optimizing the likelihood margin between preferred and dispreferred outputs under the Bradley-Terry model without anchoring to a reference, MaPO transforms diverse T2I tasks into unified pairwise preference optimization. We validate MaPO's versatility across five challenging domains: (1) safe generation, (2) style adaptation, (3) cultural representation, (4) personalization, and (5) general preference alignment. Our results reveal that MaPO's advantage grows dramatically with reference mismatch severity, outperforming both DPO and specialized methods like DreamBooth while reducing training time by 15%. MaPO thus emerges as a versatile and memory-efficient method for generic T2I adaptation tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2406_06424
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Margin-aware Preference Optimization for Aligning Diffusion Models without Reference
Hong, Jiwoo
Paul, Sayak
Lee, Noah
Rasul, Kashif
Thorne, James
Jeong, Jongheon
Computer Vision and Pattern Recognition
Modern preference alignment methods, such as DPO, rely on divergence regularization to a reference model for training stability-but this creates a fundamental problem we call "reference mismatch." In this paper, we investigate the negative impacts of reference mismatch in aligning text-to-image (T2I) diffusion models, showing that larger reference mismatch hinders effective adaptation given the same amount of data, e.g., as when learning new artistic styles, or personalizing to specific objects. We demonstrate this phenomenon across text-to-image (T2I) diffusion models and introduce margin-aware preference optimization (MaPO), a reference-agnostic approach that breaks free from this constraint. By directly optimizing the likelihood margin between preferred and dispreferred outputs under the Bradley-Terry model without anchoring to a reference, MaPO transforms diverse T2I tasks into unified pairwise preference optimization. We validate MaPO's versatility across five challenging domains: (1) safe generation, (2) style adaptation, (3) cultural representation, (4) personalization, and (5) general preference alignment. Our results reveal that MaPO's advantage grows dramatically with reference mismatch severity, outperforming both DPO and specialized methods like DreamBooth while reducing training time by 15%. MaPO thus emerges as a versatile and memory-efficient method for generic T2I adaptation tasks.
title Margin-aware Preference Optimization for Aligning Diffusion Models without Reference
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.06424