Generalizing Alignment Paradigm of Text-to-Image Generation with Preferences through $f$-divergence Minimization

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Sun, Haoyuan, Xia, Bo, Chang, Yongzhe, Wang, Xueqian
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915006831919104
author Sun, Haoyuan
Xia, Bo
Chang, Yongzhe
Wang, Xueqian
author_facet Sun, Haoyuan
Xia, Bo
Chang, Yongzhe
Wang, Xueqian
contents Direct Preference Optimization (DPO) has recently expanded its successful application from aligning large language models (LLMs) to aligning text-to-image models with human preferences, which has generated considerable interest within the community. However, we have observed that these approaches rely solely on minimizing the reverse Kullback-Leibler divergence during alignment process between the fine-tuned model and the reference model, neglecting the incorporation of other divergence constraints. In this study, we focus on extending reverse Kullback-Leibler divergence in the alignment paradigm of text-to-image models to $f$-divergence, which aims to garner better alignment performance as well as good generation diversity. We provide the generalized formula of the alignment paradigm under the $f$-divergence condition and thoroughly analyze the impact of different divergence constraints on alignment process from the perspective of gradient fields. We conduct comprehensive evaluation on image-text alignment performance, human value alignment performance and generation diversity performance under different divergence constraints, and the results indicate that alignment based on Jensen-Shannon divergence achieves the best trade-off among them. The option of divergence employed for aligning text-to-image models significantly impacts the trade-off between alignment performance (especially human value alignment) and generation diversity, which highlights the necessity of selecting an appropriate divergence for practical applications.
format Preprint
id arxiv_https___arxiv_org_abs_2409_09774
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Generalizing Alignment Paradigm of Text-to-Image Generation with Preferences through $f$-divergence Minimization
Sun, Haoyuan
Xia, Bo
Chang, Yongzhe
Wang, Xueqian
Computer Vision and Pattern Recognition
Direct Preference Optimization (DPO) has recently expanded its successful application from aligning large language models (LLMs) to aligning text-to-image models with human preferences, which has generated considerable interest within the community. However, we have observed that these approaches rely solely on minimizing the reverse Kullback-Leibler divergence during alignment process between the fine-tuned model and the reference model, neglecting the incorporation of other divergence constraints. In this study, we focus on extending reverse Kullback-Leibler divergence in the alignment paradigm of text-to-image models to $f$-divergence, which aims to garner better alignment performance as well as good generation diversity. We provide the generalized formula of the alignment paradigm under the $f$-divergence condition and thoroughly analyze the impact of different divergence constraints on alignment process from the perspective of gradient fields. We conduct comprehensive evaluation on image-text alignment performance, human value alignment performance and generation diversity performance under different divergence constraints, and the results indicate that alignment based on Jensen-Shannon divergence achieves the best trade-off among them. The option of divergence employed for aligning text-to-image models significantly impacts the trade-off between alignment performance (especially human value alignment) and generation diversity, which highlights the necessity of selecting an appropriate divergence for practical applications.
title Generalizing Alignment Paradigm of Text-to-Image Generation with Preferences through $f$-divergence Minimization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2409.09774