Half-order Fine-Tuning for Diffusion Model: A Recursive Likelihood Ratio Optimizer

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ren, Tao, Zhang, Zishi, Jiang, Jingyang, Li, Zehao, Qin, Shentao, Zheng, Yi, Li, Guanghao, Sun, Qianyou, Li, Yan, Liang, Jiafeng, Li, Xinping, Peng, Yijie
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918149382733824
author Ren, Tao
Zhang, Zishi
Jiang, Jingyang
Li, Zehao
Qin, Shentao
Zheng, Yi
Li, Guanghao
Sun, Qianyou
Li, Yan
Liang, Jiafeng
Li, Xinping
Peng, Yijie
author_facet Ren, Tao
Zhang, Zishi
Jiang, Jingyang
Li, Zehao
Qin, Shentao
Zheng, Yi
Li, Guanghao
Sun, Qianyou
Li, Yan
Liang, Jiafeng
Li, Xinping
Peng, Yijie
contents The probabilistic diffusion model (DM), generating content by inferencing through a recursive chain structure, has emerged as a powerful framework for visual generation. After pre-training on enormous data, the model needs to be properly aligned to meet requirements for downstream applications. How to efficiently align the foundation DM is a crucial task. Contemporary methods are either based on Reinforcement Learning (RL) or truncated Backpropagation (BP). However, RL and truncated BP suffer from low sample efficiency and biased gradient estimation, respectively, resulting in limited improvement or, even worse, complete training failure. To overcome the challenges, we propose the Recursive Likelihood Ratio (RLR) optimizer, a Half-Order (HO) fine-tuning paradigm for DM. The HO gradient estimator enables the computation graph rearrangement within the recursive diffusive chain, making the RLR's gradient estimator an unbiased one with lower variance than other methods. We theoretically investigate the bias, variance, and convergence of our method. Extensive experiments are conducted on image and video generation to validate the superiority of the RLR. Furthermore, we propose a novel prompt technique that is natural for the RLR to achieve a synergistic effect.
format Preprint
id arxiv_https___arxiv_org_abs_2502_00639
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Half-order Fine-Tuning for Diffusion Model: A Recursive Likelihood Ratio Optimizer
Ren, Tao
Zhang, Zishi
Jiang, Jingyang
Li, Zehao
Qin, Shentao
Zheng, Yi
Li, Guanghao
Sun, Qianyou
Li, Yan
Liang, Jiafeng
Li, Xinping
Peng, Yijie
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
The probabilistic diffusion model (DM), generating content by inferencing through a recursive chain structure, has emerged as a powerful framework for visual generation. After pre-training on enormous data, the model needs to be properly aligned to meet requirements for downstream applications. How to efficiently align the foundation DM is a crucial task. Contemporary methods are either based on Reinforcement Learning (RL) or truncated Backpropagation (BP). However, RL and truncated BP suffer from low sample efficiency and biased gradient estimation, respectively, resulting in limited improvement or, even worse, complete training failure. To overcome the challenges, we propose the Recursive Likelihood Ratio (RLR) optimizer, a Half-Order (HO) fine-tuning paradigm for DM. The HO gradient estimator enables the computation graph rearrangement within the recursive diffusive chain, making the RLR's gradient estimator an unbiased one with lower variance than other methods. We theoretically investigate the bias, variance, and convergence of our method. Extensive experiments are conducted on image and video generation to validate the superiority of the RLR. Furthermore, we propose a novel prompt technique that is natural for the RLR to achieve a synergistic effect.
title Half-order Fine-Tuning for Diffusion Model: A Recursive Likelihood Ratio Optimizer
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2502.00639