d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Pan, Leyi, Tao, Shuchang, Zhai, Yunpeng, Fu, Zheyu, Fang, Liancheng, He, Minghua, Zhang, Lingzhe, Liu, Zhaoyang, Ding, Bolin, Liu, Aiwei, Wen, Lijie
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909038723203072
author Pan, Leyi
Tao, Shuchang
Zhai, Yunpeng
Fu, Zheyu
Fang, Liancheng
He, Minghua
Zhang, Lingzhe
Liu, Zhaoyang
Ding, Bolin
Liu, Aiwei
Wen, Lijie
author_facet Pan, Leyi
Tao, Shuchang
Zhai, Yunpeng
Fu, Zheyu
Fang, Liancheng
He, Minghua
Zhang, Lingzhe
Liu, Zhaoyang
Ding, Bolin
Liu, Aiwei
Wen, Lijie
contents Reinforcement learning (RL) is pivotal for enhancing the reasoning capabilities of diffusion large language models (dLLMs). However, existing dLLM policy optimization methods suffer from two critical reliability bottlenecks: (1) reward sparsity, arising from coarse or unverifiable signals that impede accurate advantage calculation; and (2) their probability estimates do not account for the gap to the unbiased expectation over all decoding orders, which are intractable to compute. To mitigate these issues, we propose d-TreeRPO, a reliable RL framework for dLLMs that leverages tree-structured rollouts and bottom-up advantage computation based on verifiable outcome rewards to provide fine-grained and verifiable step-wise reward signals. Furthermore, we provide a theoretical proof demonstrating that increasing prediction confidence effectively minimizes the gap between unbiased expected prediction probabilities and its single-step forward pass estimate. Guided by this analysis, we introduce a time-scheduled self-distillation loss during training that enhances prediction confidence in later training stages, thereby enabling more accurate probability estimation and better performance. Experiments demonstrate that d-TreeRPO outperforms existing baselines and achieves significant improvements across multiple reasoning benchmarks. Specifically, it achieves +86.2% on Sudoku, +51.6% on Countdown, +4.5% on GSM8K, and +5.3% on Math500 compared to the base model.
format Preprint
id arxiv_https___arxiv_org_abs_2512_09675
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language Models
Pan, Leyi
Tao, Shuchang
Zhai, Yunpeng
Fu, Zheyu
Fang, Liancheng
He, Minghua
Zhang, Lingzhe
Liu, Zhaoyang
Ding, Bolin
Liu, Aiwei
Wen, Lijie
Computation and Language
68T50
I.2.7
Reinforcement learning (RL) is pivotal for enhancing the reasoning capabilities of diffusion large language models (dLLMs). However, existing dLLM policy optimization methods suffer from two critical reliability bottlenecks: (1) reward sparsity, arising from coarse or unverifiable signals that impede accurate advantage calculation; and (2) their probability estimates do not account for the gap to the unbiased expectation over all decoding orders, which are intractable to compute. To mitigate these issues, we propose d-TreeRPO, a reliable RL framework for dLLMs that leverages tree-structured rollouts and bottom-up advantage computation based on verifiable outcome rewards to provide fine-grained and verifiable step-wise reward signals. Furthermore, we provide a theoretical proof demonstrating that increasing prediction confidence effectively minimizes the gap between unbiased expected prediction probabilities and its single-step forward pass estimate. Guided by this analysis, we introduce a time-scheduled self-distillation loss during training that enhances prediction confidence in later training stages, thereby enabling more accurate probability estimation and better performance. Experiments demonstrate that d-TreeRPO outperforms existing baselines and achieves significant improvements across multiple reasoning benchmarks. Specifically, it achieves +86.2% on Sudoku, +51.6% on Countdown, +4.5% on GSM8K, and +5.3% on Math500 compared to the base model.
title d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language Models
topic Computation and Language
68T50
I.2.7
url https://arxiv.org/abs/2512.09675