InPO: Inversion Preference Optimization with Reparametrized DDIM for Efficient Diffusion Model Alignment

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lu, Yunhong, Wang, Qichao, Cao, Hengyuan, Wang, Xierui, Xu, Xiaoyin, Zhang, Min
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908280774721536
author Lu, Yunhong
Wang, Qichao
Cao, Hengyuan
Wang, Xierui
Xu, Xiaoyin
Zhang, Min
author_facet Lu, Yunhong
Wang, Qichao
Cao, Hengyuan
Wang, Xierui
Xu, Xiaoyin
Zhang, Min
contents Without using explicit reward, direct preference optimization (DPO) employs paired human preference data to fine-tune generative models, a method that has garnered considerable attention in large language models (LLMs). However, exploration of aligning text-to-image (T2I) diffusion models with human preferences remains limited. In comparison to supervised fine-tuning, existing methods that align diffusion model suffer from low training efficiency and subpar generation quality due to the long Markov chain process and the intractability of the reverse process. To address these limitations, we introduce DDIM-InPO, an efficient method for direct preference alignment of diffusion models. Our approach conceptualizes diffusion model as a single-step generative model, allowing us to fine-tune the outputs of specific latent variables selectively. In order to accomplish this objective, we first assign implicit rewards to any latent variable directly via a reparameterization technique. Then we construct an Inversion technique to estimate appropriate latent variables for preference optimization. This modification process enables the diffusion model to only fine-tune the outputs of latent variables that have a strong correlation with the preference dataset. Experimental results indicate that our DDIM-InPO achieves state-of-the-art performance with just 400 steps of fine-tuning, surpassing all preference aligning baselines for T2I diffusion models in human preference evaluation tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18454
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle InPO: Inversion Preference Optimization with Reparametrized DDIM for Efficient Diffusion Model Alignment
Lu, Yunhong
Wang, Qichao
Cao, Hengyuan
Wang, Xierui
Xu, Xiaoyin
Zhang, Min
Computer Vision and Pattern Recognition
Machine Learning
Without using explicit reward, direct preference optimization (DPO) employs paired human preference data to fine-tune generative models, a method that has garnered considerable attention in large language models (LLMs). However, exploration of aligning text-to-image (T2I) diffusion models with human preferences remains limited. In comparison to supervised fine-tuning, existing methods that align diffusion model suffer from low training efficiency and subpar generation quality due to the long Markov chain process and the intractability of the reverse process. To address these limitations, we introduce DDIM-InPO, an efficient method for direct preference alignment of diffusion models. Our approach conceptualizes diffusion model as a single-step generative model, allowing us to fine-tune the outputs of specific latent variables selectively. In order to accomplish this objective, we first assign implicit rewards to any latent variable directly via a reparameterization technique. Then we construct an Inversion technique to estimate appropriate latent variables for preference optimization. This modification process enables the diffusion model to only fine-tune the outputs of latent variables that have a strong correlation with the preference dataset. Experimental results indicate that our DDIM-InPO achieves state-of-the-art performance with just 400 steps of fine-tuning, surpassing all preference aligning baselines for T2I diffusion models in human preference evaluation tasks.
title InPO: Inversion Preference Optimization with Reparametrized DDIM for Efficient Diffusion Model Alignment
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2503.18454