Understanding Sampler Stochasticity in Training Diffusion Models for RLHF

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sheng, Jiayuan, Zhao, Hanyang, Chen, Haoxian, Yao, David D., Tang, Wenpin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914204597878784
author Sheng, Jiayuan
Zhao, Hanyang
Chen, Haoxian
Yao, David D.
Tang, Wenpin
author_facet Sheng, Jiayuan
Zhao, Hanyang
Chen, Haoxian
Yao, David D.
Tang, Wenpin
contents Reinforcement Learning from Human Feedback (RLHF) is increasingly used to fine-tune diffusion models, but a key challenge arises from the mismatch between stochastic samplers used during training and deterministic samplers used during inference. In practice, models are fine-tuned using stochastic SDE samplers to encourage exploration, while inference typically relies on deterministic ODE samplers for efficiency and stability. This discrepancy induces a reward gap, raising concerns about whether high-quality outputs can be expected during inference. In this paper, we theoretically characterize this reward gap and provide non-vacuous bounds for general diffusion models, along with sharper convergence rates for Variance Exploding (VE) and Variance Preserving (VP) Gaussian models. Methodologically, we adopt the generalized denoising diffusion implicit models (gDDIM) framework to support arbitrarily high levels of stochasticity, preserving data marginals throughout. Empirically, our findings through large-scale experiments on text-to-image models using denoising diffusion policy optimization (DDPO) and mixed group relative policy optimization (MixGRPO) validate that reward gaps consistently narrow over training, and ODE sampling quality improves when models are updated using higher-stochasticity SDE training.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10767
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Understanding Sampler Stochasticity in Training Diffusion Models for RLHF
Sheng, Jiayuan
Zhao, Hanyang
Chen, Haoxian
Yao, David D.
Tang, Wenpin
Machine Learning
Artificial Intelligence
Optimization and Control
Reinforcement Learning from Human Feedback (RLHF) is increasingly used to fine-tune diffusion models, but a key challenge arises from the mismatch between stochastic samplers used during training and deterministic samplers used during inference. In practice, models are fine-tuned using stochastic SDE samplers to encourage exploration, while inference typically relies on deterministic ODE samplers for efficiency and stability. This discrepancy induces a reward gap, raising concerns about whether high-quality outputs can be expected during inference. In this paper, we theoretically characterize this reward gap and provide non-vacuous bounds for general diffusion models, along with sharper convergence rates for Variance Exploding (VE) and Variance Preserving (VP) Gaussian models. Methodologically, we adopt the generalized denoising diffusion implicit models (gDDIM) framework to support arbitrarily high levels of stochasticity, preserving data marginals throughout. Empirically, our findings through large-scale experiments on text-to-image models using denoising diffusion policy optimization (DDPO) and mixed group relative policy optimization (MixGRPO) validate that reward gaps consistently narrow over training, and ODE sampling quality improves when models are updated using higher-stochasticity SDE training.
title Understanding Sampler Stochasticity in Training Diffusion Models for RLHF
topic Machine Learning
Artificial Intelligence
Optimization and Control
url https://arxiv.org/abs/2510.10767