A Dense Reward View on Aligning Text-to-Image Diffusion with Preference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Shentao, Chen, Tianqi, Zhou, Mingyuan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916242483314688
author Yang, Shentao
Chen, Tianqi
Zhou, Mingyuan
author_facet Yang, Shentao
Chen, Tianqi
Zhou, Mingyuan
contents Aligning text-to-image diffusion model (T2I) with preference has been gaining increasing research attention. While prior works exist on directly optimizing T2I by preference data, these methods are developed under the bandit assumption of a latent reward on the entire diffusion reverse chain, while ignoring the sequential nature of the generation process. This may harm the efficacy and efficiency of preference alignment. In this paper, we take on a finer dense reward perspective and derive a tractable alignment objective that emphasizes the initial steps of the T2I reverse chain. In particular, we introduce temporal discounting into DPO-style explicit-reward-free objectives, to break the temporal symmetry therein and suit the T2I generation hierarchy. In experiments on single and multiple prompt generation, our method is competitive with strong relevant baselines, both quantitatively and qualitatively. Further investigations are conducted to illustrate the insight of our approach.
format Preprint
id arxiv_https___arxiv_org_abs_2402_08265
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Dense Reward View on Aligning Text-to-Image Diffusion with Preference
Yang, Shentao
Chen, Tianqi
Zhou, Mingyuan
Computer Vision and Pattern Recognition
Aligning text-to-image diffusion model (T2I) with preference has been gaining increasing research attention. While prior works exist on directly optimizing T2I by preference data, these methods are developed under the bandit assumption of a latent reward on the entire diffusion reverse chain, while ignoring the sequential nature of the generation process. This may harm the efficacy and efficiency of preference alignment. In this paper, we take on a finer dense reward perspective and derive a tractable alignment objective that emphasizes the initial steps of the T2I reverse chain. In particular, we introduce temporal discounting into DPO-style explicit-reward-free objectives, to break the temporal symmetry therein and suit the T2I generation hierarchy. In experiments on single and multiple prompt generation, our method is competitive with strong relevant baselines, both quantitatively and qualitatively. Further investigations are conducted to illustrate the insight of our approach.
title A Dense Reward View on Aligning Text-to-Image Diffusion with Preference
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2402.08265