Beyond Pairwise Preferences: Listwise Reward-Aware Alignment for Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Austin, Han, Jiaqi, Ermon, Stefano, Yue, Yisong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914602552393728
author Wang, Austin
Han, Jiaqi
Ermon, Stefano
Yue, Yisong
author_facet Wang, Austin
Han, Jiaqi
Ermon, Stefano
Yue, Yisong
contents Preference optimization has emerged as an efficient alternative to online reinforcement learning from human feedback (RLHF) for aligning text-to-image diffusion models. However, existing methods largely reduce supervision to binary pairwise comparisons. This pairwise reduction is limiting when training data naturally contains multiple candidate images for the same prompt, and when continuous reward scores can provide richer information than a single winner-loser label. To address these limitations, we propose Diffusion LAIR, a reward-aware listwise preference optimization method for diffusion models. For each prompt, LAIR converts reward scores across a group of candidate images into centered advantage weights, then optimizes an advantage-weighted regression objective on the implicit reward, defined as the denoising-loss improvement of the current model over a fixed reference model, with a quadratic penalty that regularizes the magnitude of the implicit reward. The resulting objective uses all candidates simultaneously rather than selecting pairs, and remains conservative by explicitly controlling the magnitude of the implicit reward. The LAIR objective admits a bounded closed-form optimum in implicit-reward space, clarifying how the regularization strength controls the magnitude of the preference update. Experiments show that Diffusion LAIR outperforms strong preference optimization baselines on SD1.5 and SDXL across text-to-image generation, compositional generation, and image editing benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2605_26491
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Beyond Pairwise Preferences: Listwise Reward-Aware Alignment for Diffusion Models
Wang, Austin
Han, Jiaqi
Ermon, Stefano
Yue, Yisong
Machine Learning
Computer Vision and Pattern Recognition
Preference optimization has emerged as an efficient alternative to online reinforcement learning from human feedback (RLHF) for aligning text-to-image diffusion models. However, existing methods largely reduce supervision to binary pairwise comparisons. This pairwise reduction is limiting when training data naturally contains multiple candidate images for the same prompt, and when continuous reward scores can provide richer information than a single winner-loser label. To address these limitations, we propose Diffusion LAIR, a reward-aware listwise preference optimization method for diffusion models. For each prompt, LAIR converts reward scores across a group of candidate images into centered advantage weights, then optimizes an advantage-weighted regression objective on the implicit reward, defined as the denoising-loss improvement of the current model over a fixed reference model, with a quadratic penalty that regularizes the magnitude of the implicit reward. The resulting objective uses all candidates simultaneously rather than selecting pairs, and remains conservative by explicitly controlling the magnitude of the implicit reward. The LAIR objective admits a bounded closed-form optimum in implicit-reward space, clarifying how the regularization strength controls the magnitude of the preference update. Experiments show that Diffusion LAIR outperforms strong preference optimization baselines on SD1.5 and SDXL across text-to-image generation, compositional generation, and image editing benchmarks.
title Beyond Pairwise Preferences: Listwise Reward-Aware Alignment for Diffusion Models
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.26491