Diffusion Model as a Noise-Aware Latent Reward Model for Step-Level Preference Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Tao, Da, Cheng, Ding, Kun, Yang, Huan, Jin, Kun, Li, Yan, Gao, Tingting, Zhang, Di, Xiang, Shiming, Pan, Chunhong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914070956867584
author Zhang, Tao
Da, Cheng
Ding, Kun
Yang, Huan
Jin, Kun
Li, Yan
Gao, Tingting
Zhang, Di
Xiang, Shiming
Pan, Chunhong
author_facet Zhang, Tao
Da, Cheng
Ding, Kun
Yang, Huan
Jin, Kun
Li, Yan
Gao, Tingting
Zhang, Di
Xiang, Shiming
Pan, Chunhong
contents Preference optimization for diffusion models aims to align them with human preferences for images. Previous methods typically use Vision-Language Models (VLMs) as pixel-level reward models to approximate human preferences. However, when used for step-level preference optimization, these models face challenges in handling noisy images of different timesteps and require complex transformations into pixel space. In this work, we show that pre-trained diffusion models are naturally suited for step-level reward modeling in the noisy latent space, as they are explicitly designed to process latent images at various noise levels. Accordingly, we propose the Latent Reward Model (LRM), which repurposes components of the diffusion model to predict preferences of latent images at arbitrary timesteps. Building on LRM, we introduce Latent Preference Optimization (LPO), a step-level preference optimization method conducted directly in the noisy latent space. Experimental results indicate that LPO significantly improves the model's alignment with general, aesthetic, and text-image alignment preferences, while achieving a 2.5-28x training speedup over existing preference optimization methods. Our code and models are available at https://github.com/Kwai-Kolors/LPO.
format Preprint
id arxiv_https___arxiv_org_abs_2502_01051
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Diffusion Model as a Noise-Aware Latent Reward Model for Step-Level Preference Optimization
Zhang, Tao
Da, Cheng
Ding, Kun
Yang, Huan
Jin, Kun
Li, Yan
Gao, Tingting
Zhang, Di
Xiang, Shiming
Pan, Chunhong
Computer Vision and Pattern Recognition
Preference optimization for diffusion models aims to align them with human preferences for images. Previous methods typically use Vision-Language Models (VLMs) as pixel-level reward models to approximate human preferences. However, when used for step-level preference optimization, these models face challenges in handling noisy images of different timesteps and require complex transformations into pixel space. In this work, we show that pre-trained diffusion models are naturally suited for step-level reward modeling in the noisy latent space, as they are explicitly designed to process latent images at various noise levels. Accordingly, we propose the Latent Reward Model (LRM), which repurposes components of the diffusion model to predict preferences of latent images at arbitrary timesteps. Building on LRM, we introduce Latent Preference Optimization (LPO), a step-level preference optimization method conducted directly in the noisy latent space. Experimental results indicate that LPO significantly improves the model's alignment with general, aesthetic, and text-image alignment preferences, while achieving a 2.5-28x training speedup over existing preference optimization methods. Our code and models are available at https://github.com/Kwai-Kolors/LPO.
title Diffusion Model as a Noise-Aware Latent Reward Model for Step-Level Preference Optimization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.01051