Aligning Few-Step Diffusion Models with Dense Reward Difference Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Ziyi, Shen, Li, Zhang, Sen, Ye, Deheng, Luo, Yong, Shi, Miaojing, Shan, Dongjing, Du, Bo, Tao, Dacheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914354438340608
author Zhang, Ziyi
Shen, Li
Zhang, Sen
Ye, Deheng
Luo, Yong
Shi, Miaojing
Shan, Dongjing
Du, Bo
Tao, Dacheng
author_facet Zhang, Ziyi
Shen, Li
Zhang, Sen
Ye, Deheng
Luo, Yong
Shi, Miaojing
Shan, Dongjing
Du, Bo
Tao, Dacheng
contents Few-step diffusion models enable efficient high-resolution image synthesis but struggle to align with specific downstream objectives due to limitations of existing reinforcement learning (RL) methods in low-step regimes with limited state spaces and suboptimal sample quality. To address this, we propose Stepwise Diffusion Policy Optimization (SDPO), a novel RL framework tailored for few-step diffusion models. SDPO introduces a dual-state trajectory sampling mechanism, tracking both noisy and predicted clean states at each step to provide dense reward feedback and enable low-variance, mixed-step optimization. For further efficiency, we develop a latent similarity-based dense reward prediction strategy to minimize costly dense reward queries. Leveraging these dense rewards, SDPO optimizes a dense reward difference learning objective that enables more frequent and granular policy updates. Additional refinements, including stepwise advantage estimates, temporal importance weighting, and step-shuffled gradient updates, further enhance long-term dependency, low-step priority, and gradient stability. Our experiments demonstrate that SDPO consistently delivers superior reward-aligned results across diverse few-step settings and tasks. Code is available at https://github.com/ZiyiZhang27/sdpo.
format Preprint
id arxiv_https___arxiv_org_abs_2411_11727
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Aligning Few-Step Diffusion Models with Dense Reward Difference Learning
Zhang, Ziyi
Shen, Li
Zhang, Sen
Ye, Deheng
Luo, Yong
Shi, Miaojing
Shan, Dongjing
Du, Bo
Tao, Dacheng
Machine Learning
Computer Vision and Pattern Recognition
Few-step diffusion models enable efficient high-resolution image synthesis but struggle to align with specific downstream objectives due to limitations of existing reinforcement learning (RL) methods in low-step regimes with limited state spaces and suboptimal sample quality. To address this, we propose Stepwise Diffusion Policy Optimization (SDPO), a novel RL framework tailored for few-step diffusion models. SDPO introduces a dual-state trajectory sampling mechanism, tracking both noisy and predicted clean states at each step to provide dense reward feedback and enable low-variance, mixed-step optimization. For further efficiency, we develop a latent similarity-based dense reward prediction strategy to minimize costly dense reward queries. Leveraging these dense rewards, SDPO optimizes a dense reward difference learning objective that enables more frequent and granular policy updates. Additional refinements, including stepwise advantage estimates, temporal importance weighting, and step-shuffled gradient updates, further enhance long-term dependency, low-step priority, and gradient stability. Our experiments demonstrate that SDPO consistently delivers superior reward-aligned results across diverse few-step settings and tasks. Code is available at https://github.com/ZiyiZhang27/sdpo.
title Aligning Few-Step Diffusion Models with Dense Reward Difference Learning
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.11727