Enhancing Reasoning for Diffusion LLMs via Distribution Matching Policy Optimization

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhu, Yuchen, Guo, Wei, Choi, Jaemoo, Molodyk, Petr, Yuan, Bo, Tao, Molei, Chen, Yongxin
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918350018314240
author Zhu, Yuchen
Guo, Wei
Choi, Jaemoo
Molodyk, Petr
Yuan, Bo
Tao, Molei
Chen, Yongxin
author_facet Zhu, Yuchen
Guo, Wei
Choi, Jaemoo
Molodyk, Petr
Yuan, Bo
Tao, Molei
Chen, Yongxin
contents Diffusion large language models (dLLMs) are promising alternatives to autoregressive large language models (AR-LLMs), as they potentially allow higher inference throughput. Reinforcement learning (RL) is a crucial component for dLLMs to achieve comparable performance with AR-LLMs on important tasks, such as reasoning. However, RL algorithms that are well-suited for dLLMs' unique characteristics have yet to be developed. This paper proposes Distribution Matching Policy Optimization (DMPO), a principled and theoretically grounded RL fine-tuning method specifically designed to enhance the reasoning capabilities of dLLMs by matching the dLLM policy distribution to the optimal, reward-tilted one through cross-entropy optimization. We identify a key challenge in the implementation with a small training batch size and propose several effective solutions through a novel weight baseline subtraction technique. DMPO exhibits superior performance on multiple reasoning benchmarks without supervised fine-tuning, with an accuracy improvement of up to $54.3\%$ over previously SOTA baselines and $66.41\%$ over the base model, underscoring the effectiveness of the distribution matching framework. Our code is available at https://github.com/yuchen-zhu-zyc/DMPO.
format Preprint
id arxiv_https___arxiv_org_abs_2510_08233
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing Reasoning for Diffusion LLMs via Distribution Matching Policy Optimization
Zhu, Yuchen
Guo, Wei
Choi, Jaemoo
Molodyk, Petr
Yuan, Bo
Tao, Molei
Chen, Yongxin
Machine Learning
Diffusion large language models (dLLMs) are promising alternatives to autoregressive large language models (AR-LLMs), as they potentially allow higher inference throughput. Reinforcement learning (RL) is a crucial component for dLLMs to achieve comparable performance with AR-LLMs on important tasks, such as reasoning. However, RL algorithms that are well-suited for dLLMs' unique characteristics have yet to be developed. This paper proposes Distribution Matching Policy Optimization (DMPO), a principled and theoretically grounded RL fine-tuning method specifically designed to enhance the reasoning capabilities of dLLMs by matching the dLLM policy distribution to the optimal, reward-tilted one through cross-entropy optimization. We identify a key challenge in the implementation with a small training batch size and propose several effective solutions through a novel weight baseline subtraction technique. DMPO exhibits superior performance on multiple reasoning benchmarks without supervised fine-tuning, with an accuracy improvement of up to $54.3\%$ over previously SOTA baselines and $66.41\%$ over the base model, underscoring the effectiveness of the distribution matching framework. Our code is available at https://github.com/yuchen-zhu-zyc/DMPO.
title Enhancing Reasoning for Diffusion LLMs via Distribution Matching Policy Optimization
topic Machine Learning
url https://arxiv.org/abs/2510.08233