Towards General Preference Alignment: Diffusion Models at Nash Equilibrium

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hu, Jiaming, Bai, Jiamu, Wang, Haoyu, Mukherjee, Debarghya, Paschalidis, Ioannis Ch.
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913094019579904
author Hu, Jiaming
Bai, Jiamu
Wang, Haoyu
Mukherjee, Debarghya
Paschalidis, Ioannis Ch.
author_facet Hu, Jiaming
Bai, Jiamu
Wang, Haoyu
Mukherjee, Debarghya
Paschalidis, Ioannis Ch.
contents Reinforcement learning from human feedback (RLHF) has been popular for aligning text-to-image (T2I) diffusion models with human preferences. As a mainstream branch of RLHF, Direct Preference Optimization (DPO) offers a computationally efficient alternative that avoids explicit reward modeling and has been widely adopted in diffusion alignment. However, existing preference-based methods for diffusion alignment still rely on reward-induced preference signals and typically assume that human preferences can be adequately modeled by the Bradley--Terry (BT) model, which may fail to capture the full complexity of human preferences. In this paper, we formulate diffusion alignment from a game-theoretic perspective. We propose Diffusion Nash Preference Optimization (Diff.-NPO), an intuitive general preference framework for diffusion alignment. Diff.-NPO encourages the current policy to play against itself to achieve self improvement and lead to a better alignment. Empirically, we demonstrate the effectiveness of Diff.-NPO on the text-to-image generation task via various metrics. Diff.-NPO consistently outperforms existing preference-based diffusion alignment methods.
format Preprint
id arxiv_https___arxiv_org_abs_2605_04494
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Towards General Preference Alignment: Diffusion Models at Nash Equilibrium
Hu, Jiaming
Bai, Jiamu
Wang, Haoyu
Mukherjee, Debarghya
Paschalidis, Ioannis Ch.
Machine Learning
Computer Vision and Pattern Recognition
Reinforcement learning from human feedback (RLHF) has been popular for aligning text-to-image (T2I) diffusion models with human preferences. As a mainstream branch of RLHF, Direct Preference Optimization (DPO) offers a computationally efficient alternative that avoids explicit reward modeling and has been widely adopted in diffusion alignment. However, existing preference-based methods for diffusion alignment still rely on reward-induced preference signals and typically assume that human preferences can be adequately modeled by the Bradley--Terry (BT) model, which may fail to capture the full complexity of human preferences. In this paper, we formulate diffusion alignment from a game-theoretic perspective. We propose Diffusion Nash Preference Optimization (Diff.-NPO), an intuitive general preference framework for diffusion alignment. Diff.-NPO encourages the current policy to play against itself to achieve self improvement and lead to a better alignment. Empirically, we demonstrate the effectiveness of Diff.-NPO on the text-to-image generation task via various metrics. Diff.-NPO consistently outperforms existing preference-based diffusion alignment methods.
title Towards General Preference Alignment: Diffusion Models at Nash Equilibrium
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.04494