Training Diffusion Models with Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Black, Kevin, Janner, Michael, Du, Yilun, Kostrikov, Ilya, Levine, Sergey
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916081168285696
author Black, Kevin
Janner, Michael
Du, Yilun
Kostrikov, Ilya
Levine, Sergey
author_facet Black, Kevin
Janner, Michael
Du, Yilun
Kostrikov, Ilya
Levine, Sergey
contents Diffusion models are a class of flexible generative models trained with an approximation to the log-likelihood objective. However, most use cases of diffusion models are not concerned with likelihoods, but instead with downstream objectives such as human-perceived image quality or drug effectiveness. In this paper, we investigate reinforcement learning methods for directly optimizing diffusion models for such objectives. We describe how posing denoising as a multi-step decision-making problem enables a class of policy gradient algorithms, which we refer to as denoising diffusion policy optimization (DDPO), that are more effective than alternative reward-weighted likelihood approaches. Empirically, DDPO is able to adapt text-to-image diffusion models to objectives that are difficult to express via prompting, such as image compressibility, and those derived from human feedback, such as aesthetic quality. Finally, we show that DDPO can improve prompt-image alignment using feedback from a vision-language model without the need for additional data collection or human annotation. The project's website can be found at http://rl-diffusion.github.io .
format Preprint
id arxiv_https___arxiv_org_abs_2305_13301
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Training Diffusion Models with Reinforcement Learning
Black, Kevin
Janner, Michael
Du, Yilun
Kostrikov, Ilya
Levine, Sergey
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Diffusion models are a class of flexible generative models trained with an approximation to the log-likelihood objective. However, most use cases of diffusion models are not concerned with likelihoods, but instead with downstream objectives such as human-perceived image quality or drug effectiveness. In this paper, we investigate reinforcement learning methods for directly optimizing diffusion models for such objectives. We describe how posing denoising as a multi-step decision-making problem enables a class of policy gradient algorithms, which we refer to as denoising diffusion policy optimization (DDPO), that are more effective than alternative reward-weighted likelihood approaches. Empirically, DDPO is able to adapt text-to-image diffusion models to objectives that are difficult to express via prompting, such as image compressibility, and those derived from human feedback, such as aesthetic quality. Finally, we show that DDPO can improve prompt-image alignment using feedback from a vision-language model without the need for additional data collection or human annotation. The project's website can be found at http://rl-diffusion.github.io .
title Training Diffusion Models with Reinforcement Learning
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2305.13301