VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Dai, Fengyuan, Zhuang, Zifeng, Huang, Yufei, Huang, Siteng, Liao, Bangyan, Wang, Donglin, Yuan, Fajie
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916773146656768
author Dai, Fengyuan
Zhuang, Zifeng
Huang, Yufei
Huang, Siteng
Liao, Bangyan
Wang, Donglin
Yuan, Fajie
author_facet Dai, Fengyuan
Zhuang, Zifeng
Huang, Yufei
Huang, Siteng
Liao, Bangyan
Wang, Donglin
Yuan, Fajie
contents Diffusion models have emerged as powerful generative tools across various domains, yet tailoring pre-trained models to exhibit specific desirable properties remains challenging. While reinforcement learning (RL) offers a promising solution,current methods struggle to simultaneously achieve stable, efficient fine-tuning and support non-differentiable rewards. Furthermore, their reliance on sparse rewards provides inadequate supervision during intermediate steps, often resulting in suboptimal generation quality. To address these limitations, dense and differentiable signals are required throughout the diffusion process. Hence, we propose VAlue-based Reinforced Diffusion (VARD): a novel approach that first learns a value function predicting expection of rewards from intermediate states, and subsequently uses this value function with KL regularization to provide dense supervision throughout the generation process. Our method maintains proximity to the pretrained model while enabling effective and stable training via backpropagation. Experimental results demonstrate that our approach facilitates better trajectory guidance, improves training efficiency and extends the applicability of RL to diffusion models optimized for complex, non-differentiable reward functions.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15791
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL
Dai, Fengyuan
Zhuang, Zifeng
Huang, Yufei
Huang, Siteng
Liao, Bangyan
Wang, Donglin
Yuan, Fajie
Computer Vision and Pattern Recognition
Machine Learning
Diffusion models have emerged as powerful generative tools across various domains, yet tailoring pre-trained models to exhibit specific desirable properties remains challenging. While reinforcement learning (RL) offers a promising solution,current methods struggle to simultaneously achieve stable, efficient fine-tuning and support non-differentiable rewards. Furthermore, their reliance on sparse rewards provides inadequate supervision during intermediate steps, often resulting in suboptimal generation quality. To address these limitations, dense and differentiable signals are required throughout the diffusion process. Hence, we propose VAlue-based Reinforced Diffusion (VARD): a novel approach that first learns a value function predicting expection of rewards from intermediate states, and subsequently uses this value function with KL regularization to provide dense supervision throughout the generation process. Our method maintains proximity to the pretrained model while enabling effective and stable training via backpropagation. Experimental results demonstrate that our approach facilitates better trajectory guidance, improves training efficiency and extends the applicability of RL to diffusion models optimized for complex, non-differentiable reward functions.
title VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2505.15791