DiffusionNFT: Online Diffusion Reinforcement with Forward Process

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Kaiwen, Chen, Huayu, Ye, Haotian, Wang, Haoxiang, Zhang, Qinsheng, Jiang, Kai, Su, Hang, Ermon, Stefano, Zhu, Jun, Liu, Ming-Yu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911450492043264
author Zheng, Kaiwen
Chen, Huayu
Ye, Haotian
Wang, Haoxiang
Zhang, Qinsheng
Jiang, Kai
Su, Hang
Ermon, Stefano
Zhu, Jun
Liu, Ming-Yu
author_facet Zheng, Kaiwen
Chen, Huayu
Ye, Haotian
Wang, Haoxiang
Zhang, Qinsheng
Jiang, Kai
Su, Hang
Ermon, Stefano
Zhu, Jun
Liu, Ming-Yu
contents Online reinforcement learning (RL) has been central to post-training language models, but its extension to diffusion models remains challenging due to intractable likelihoods. Recent works discretize the reverse sampling process to enable GRPO-style training, yet they inherit fundamental drawbacks, including solver restrictions, forward-reverse inconsistency, and complicated integration with classifier-free guidance (CFG). We introduce Diffusion Negative-aware FineTuning (DiffusionNFT), a new online RL paradigm that optimizes diffusion models directly on the forward process via flow matching. DiffusionNFT contrasts positive and negative generations to define an implicit policy improvement direction, naturally incorporating reinforcement signals into the supervised learning objective. This formulation enables training with arbitrary black-box solvers, eliminates the need for likelihood estimation, and requires only clean images rather than sampling trajectories for policy optimization. DiffusionNFT is up to $25\times$ more efficient than FlowGRPO in head-to-head comparisons, while being CFG-free. For instance, DiffusionNFT improves the GenEval score from 0.24 to 0.98 within 1k steps, while FlowGRPO achieves 0.95 with over 5k steps and additional CFG employment. By leveraging multiple reward models, DiffusionNFT significantly boosts the performance of SD3.5-Medium in every benchmark tested.
format Preprint
id arxiv_https___arxiv_org_abs_2509_16117
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DiffusionNFT: Online Diffusion Reinforcement with Forward Process
Zheng, Kaiwen
Chen, Huayu
Ye, Haotian
Wang, Haoxiang
Zhang, Qinsheng
Jiang, Kai
Su, Hang
Ermon, Stefano
Zhu, Jun
Liu, Ming-Yu
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Online reinforcement learning (RL) has been central to post-training language models, but its extension to diffusion models remains challenging due to intractable likelihoods. Recent works discretize the reverse sampling process to enable GRPO-style training, yet they inherit fundamental drawbacks, including solver restrictions, forward-reverse inconsistency, and complicated integration with classifier-free guidance (CFG). We introduce Diffusion Negative-aware FineTuning (DiffusionNFT), a new online RL paradigm that optimizes diffusion models directly on the forward process via flow matching. DiffusionNFT contrasts positive and negative generations to define an implicit policy improvement direction, naturally incorporating reinforcement signals into the supervised learning objective. This formulation enables training with arbitrary black-box solvers, eliminates the need for likelihood estimation, and requires only clean images rather than sampling trajectories for policy optimization. DiffusionNFT is up to $25\times$ more efficient than FlowGRPO in head-to-head comparisons, while being CFG-free. For instance, DiffusionNFT improves the GenEval score from 0.24 to 0.98 within 1k steps, while FlowGRPO achieves 0.95 with over 5k steps and additional CFG employment. By leveraging multiple reward models, DiffusionNFT significantly boosts the performance of SD3.5-Medium in every benchmark tested.
title DiffusionNFT: Online Diffusion Reinforcement with Forward Process
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.16117