Linear-DPO: Linear Direct Preference Optimization for Diffusion and Flow-Matching Generative Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Kesong, Xu, Yixuan, Tseng, Kuo-kun, Lu, Weiyi, Liu, Kan, Lan, Tao
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914583152689152
author Li, Kesong
Xu, Yixuan
Tseng, Kuo-kun
Lu, Weiyi
Liu, Kan
Lan, Tao
author_facet Li, Kesong
Xu, Yixuan
Tseng, Kuo-kun
Lu, Weiyi
Liu, Kan
Lan, Tao
contents Direct Preference Optimization (DPO) is successful for alignment in LLMs but still faces challenges in text-to-image generation. Existing studies are confined to denoising diffusion models while overlooking flow-matching, and suffer from an objective mismatch when applying discrete NLP-based DPO to regression-based generative tasks.\ In this paper, we derive a generalized DPO objective that covers both diffusion and flow-matching via a unified reverse-time SDE framework, and point out from a gradient perspective that the standard DPO objective is suboptimal for text-to-image generation. Consequently, we propose Linear-DPO, which replaces the aggressive sigmoid-based utility function with a sustained linear utility and incorporates an EMA-updated reference model. Qualitative and quantitative experiments on diffusion models (SD1.5, SDXL) and flow-matching model (SD3-Medium) demonstrate the superiority of our approach over existing baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2605_21123
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Linear-DPO: Linear Direct Preference Optimization for Diffusion and Flow-Matching Generative Models
Li, Kesong
Xu, Yixuan
Tseng, Kuo-kun
Lu, Weiyi
Liu, Kan
Lan, Tao
Computer Vision and Pattern Recognition
Machine Learning
I.2.6; I.4.9; G.3
Direct Preference Optimization (DPO) is successful for alignment in LLMs but still faces challenges in text-to-image generation. Existing studies are confined to denoising diffusion models while overlooking flow-matching, and suffer from an objective mismatch when applying discrete NLP-based DPO to regression-based generative tasks.\ In this paper, we derive a generalized DPO objective that covers both diffusion and flow-matching via a unified reverse-time SDE framework, and point out from a gradient perspective that the standard DPO objective is suboptimal for text-to-image generation. Consequently, we propose Linear-DPO, which replaces the aggressive sigmoid-based utility function with a sustained linear utility and incorporates an EMA-updated reference model. Qualitative and quantitative experiments on diffusion models (SD1.5, SDXL) and flow-matching model (SD3-Medium) demonstrate the superiority of our approach over existing baselines.
title Linear-DPO: Linear Direct Preference Optimization for Diffusion and Flow-Matching Generative Models
topic Computer Vision and Pattern Recognition
Machine Learning
I.2.6; I.4.9; G.3
url https://arxiv.org/abs/2605.21123