Rethinking the Design Space of Reinforcement Learning for Diffusion Models: On the Importance of Likelihood Estimation Beyond Loss Design

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Choi, Jaemoo, Zhu, Yuchen, Guo, Wei, Molodyk, Petr, Yuan, Bo, Bai, Jinbin, Xin, Yi, Tao, Molei, Chen, Yongxin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911695566274560
author Choi, Jaemoo
Zhu, Yuchen
Guo, Wei
Molodyk, Petr
Yuan, Bo
Bai, Jinbin
Xin, Yi
Tao, Molei
Chen, Yongxin
author_facet Choi, Jaemoo
Zhu, Yuchen
Guo, Wei
Molodyk, Petr
Yuan, Bo
Bai, Jinbin
Xin, Yi
Tao, Molei
Chen, Yongxin
contents Reinforcement learning has been widely applied to diffusion and flow models for visual tasks such as text-to-image generation. However, these tasks remain challenging because diffusion models have intractable likelihoods, which creates a barrier for directly applying popular policy-gradient type methods. Existing approaches primarily focus on crafting new objectives built on already heavily engineered LLM objectives, using ad hoc estimators for likelihood, without a thorough investigation into how such estimation affects overall algorithmic performance. In this work, we provide a systematic analysis of the RL design space by disentangling three factors: i) policy-gradient objectives, ii) likelihood estimators, and iii) rollout sampling schemes. We show that adopting an evidence lower bound (ELBO) based model likelihood estimator, computed only from the final generated sample, is the dominant factor enabling effective, efficient, and stable RL optimization, outweighing the impact of the specific policy-gradient loss functional. We validate our findings across multiple reward benchmarks using SD 3.5 Medium, and observe consistent trends across all tasks. Our method improves the GenEval score from 0.24 to 0.95 in 90 GPU hours, which is $4.6\times$ more efficient than FlowGRPO and $2\times$ more efficient than the SOTA method DiffusionNFT without reward hacking.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04663
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Rethinking the Design Space of Reinforcement Learning for Diffusion Models: On the Importance of Likelihood Estimation Beyond Loss Design
Choi, Jaemoo
Zhu, Yuchen
Guo, Wei
Molodyk, Petr
Yuan, Bo
Bai, Jinbin
Xin, Yi
Tao, Molei
Chen, Yongxin
Machine Learning
Artificial Intelligence
Reinforcement learning has been widely applied to diffusion and flow models for visual tasks such as text-to-image generation. However, these tasks remain challenging because diffusion models have intractable likelihoods, which creates a barrier for directly applying popular policy-gradient type methods. Existing approaches primarily focus on crafting new objectives built on already heavily engineered LLM objectives, using ad hoc estimators for likelihood, without a thorough investigation into how such estimation affects overall algorithmic performance. In this work, we provide a systematic analysis of the RL design space by disentangling three factors: i) policy-gradient objectives, ii) likelihood estimators, and iii) rollout sampling schemes. We show that adopting an evidence lower bound (ELBO) based model likelihood estimator, computed only from the final generated sample, is the dominant factor enabling effective, efficient, and stable RL optimization, outweighing the impact of the specific policy-gradient loss functional. We validate our findings across multiple reward benchmarks using SD 3.5 Medium, and observe consistent trends across all tasks. Our method improves the GenEval score from 0.24 to 0.95 in 90 GPU hours, which is $4.6\times$ more efficient than FlowGRPO and $2\times$ more efficient than the SOTA method DiffusionNFT without reward hacking.
title Rethinking the Design Space of Reinforcement Learning for Diffusion Models: On the Importance of Likelihood Estimation Beyond Loss Design
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.04663