Consolidating Reinforcement Learning for Multimodal Discrete Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Tianren, Zhang, Mu, Wang, Yibing, Ye, Qixiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914072604180480
author Ma, Tianren
Zhang, Mu
Wang, Yibing
Ye, Qixiang
author_facet Ma, Tianren
Zhang, Mu
Wang, Yibing
Ye, Qixiang
contents Optimizing discrete diffusion model (DDM) with rewards remains a challenge: the non-autoregressive paradigm makes importance sampling intractable and rollout complex, puzzling reinforcement learning methods such as Group Relative Policy Optimization (GRPO). In this study, we introduce MaskGRPO, the first viable approach to enable scalable multimodal reinforcement learning in discrete diffusion with effective importance sampling and modality-specific adaptations. To this end, we first clarify the theoretical foundation for DDMs, which facilitates building an importance estimator that captures valuable token fluctuation for gradient updates. We then delicately tailored the rollout method for visual sequences, which yields diverse completions and reliable optimization gradients. Upon math reasoning, coding, and visual generation benchmarks, MaskGRPO brings more stable and efficient updates, leading to stronger reasoning performance and better generation quality. This study establishes MaskGRPO as a systematic policy optimization approach and the first practical way for discretized visual diffusion.
format Preprint
id arxiv_https___arxiv_org_abs_2510_02880
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Consolidating Reinforcement Learning for Multimodal Discrete Diffusion Models
Ma, Tianren
Zhang, Mu
Wang, Yibing
Ye, Qixiang
Artificial Intelligence
Optimizing discrete diffusion model (DDM) with rewards remains a challenge: the non-autoregressive paradigm makes importance sampling intractable and rollout complex, puzzling reinforcement learning methods such as Group Relative Policy Optimization (GRPO). In this study, we introduce MaskGRPO, the first viable approach to enable scalable multimodal reinforcement learning in discrete diffusion with effective importance sampling and modality-specific adaptations. To this end, we first clarify the theoretical foundation for DDMs, which facilitates building an importance estimator that captures valuable token fluctuation for gradient updates. We then delicately tailored the rollout method for visual sequences, which yields diverse completions and reliable optimization gradients. Upon math reasoning, coding, and visual generation benchmarks, MaskGRPO brings more stable and efficient updates, leading to stronger reasoning performance and better generation quality. This study establishes MaskGRPO as a systematic policy optimization approach and the first practical way for discretized visual diffusion.
title Consolidating Reinforcement Learning for Multimodal Discrete Diffusion Models
topic Artificial Intelligence
url https://arxiv.org/abs/2510.02880