Linear-DPO: Linear Direct Preference Optimization for Diffusion and Flow-Matching Generative Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866914583152689152 |
|---|---|
| author | Li, Kesong Xu, Yixuan Tseng, Kuo-kun Lu, Weiyi Liu, Kan Lan, Tao |
| author_facet | Li, Kesong Xu, Yixuan Tseng, Kuo-kun Lu, Weiyi Liu, Kan Lan, Tao |
| contents | Direct Preference Optimization (DPO) is successful for alignment in LLMs but still faces challenges in text-to-image generation. Existing studies are confined to denoising diffusion models while overlooking flow-matching, and suffer from an objective mismatch when applying discrete NLP-based DPO to regression-based generative tasks.\ In this paper, we derive a generalized DPO objective that covers both diffusion and flow-matching via a unified reverse-time SDE framework, and point out from a gradient perspective that the standard DPO objective is suboptimal for text-to-image generation. Consequently, we propose Linear-DPO, which replaces the aggressive sigmoid-based utility function with a sustained linear utility and incorporates an EMA-updated reference model. Qualitative and quantitative experiments on diffusion models (SD1.5, SDXL) and flow-matching model (SD3-Medium) demonstrate the superiority of our approach over existing baselines. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_21123 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Linear-DPO: Linear Direct Preference Optimization for Diffusion and Flow-Matching Generative Models Li, Kesong Xu, Yixuan Tseng, Kuo-kun Lu, Weiyi Liu, Kan Lan, Tao Computer Vision and Pattern Recognition Machine Learning I.2.6; I.4.9; G.3 Direct Preference Optimization (DPO) is successful for alignment in LLMs but still faces challenges in text-to-image generation. Existing studies are confined to denoising diffusion models while overlooking flow-matching, and suffer from an objective mismatch when applying discrete NLP-based DPO to regression-based generative tasks.\ In this paper, we derive a generalized DPO objective that covers both diffusion and flow-matching via a unified reverse-time SDE framework, and point out from a gradient perspective that the standard DPO objective is suboptimal for text-to-image generation. Consequently, we propose Linear-DPO, which replaces the aggressive sigmoid-based utility function with a sustained linear utility and incorporates an EMA-updated reference model. Qualitative and quantitative experiments on diffusion models (SD1.5, SDXL) and flow-matching model (SD3-Medium) demonstrate the superiority of our approach over existing baselines. |
| title | Linear-DPO: Linear Direct Preference Optimization for Diffusion and Flow-Matching Generative Models |
| topic | Computer Vision and Pattern Recognition Machine Learning I.2.6; I.4.9; G.3 |
| url | https://arxiv.org/abs/2605.21123 |