Fine-Tuning Discrete Diffusion Models with Policy Gradient Methods

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zekri, Oussama, Boullé, Nicolas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912771794272256
author Zekri, Oussama
Boullé, Nicolas
author_facet Zekri, Oussama
Boullé, Nicolas
contents Discrete diffusion models have recently gained significant attention due to their ability to process complex discrete structures for language modeling. However, fine-tuning these models with policy gradient methods, as is commonly done in Reinforcement Learning from Human Feedback (RLHF), remains a challenging task. We propose an efficient, broadly applicable, and theoretically justified policy gradient algorithm, called Score Entropy Policy Optimization (\SEPO), for fine-tuning discrete diffusion models over non-differentiable rewards. Our numerical experiments across several discrete generative tasks demonstrate the scalability and efficiency of our method. Our code is available at https://github.com/ozekri/SEPO.
format Preprint
id arxiv_https___arxiv_org_abs_2502_01384
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fine-Tuning Discrete Diffusion Models with Policy Gradient Methods
Zekri, Oussama
Boullé, Nicolas
Machine Learning
Artificial Intelligence
Computation and Language
Discrete diffusion models have recently gained significant attention due to their ability to process complex discrete structures for language modeling. However, fine-tuning these models with policy gradient methods, as is commonly done in Reinforcement Learning from Human Feedback (RLHF), remains a challenging task. We propose an efficient, broadly applicable, and theoretically justified policy gradient algorithm, called Score Entropy Policy Optimization (\SEPO), for fine-tuning discrete diffusion models over non-differentiable rewards. Our numerical experiments across several discrete generative tasks demonstrate the scalability and efficiency of our method. Our code is available at https://github.com/ozekri/SEPO.
title Fine-Tuning Discrete Diffusion Models with Policy Gradient Methods
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2502.01384