The Flexibility Trap: Why Arbitrary Order Limits Reasoning Potential in Diffusion Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ni, Zanlin, Wang, Shenzhi, Yue, Yang, Yu, Tianyu, Zhao, Weilin, Hua, Yeguo, Chen, Tianyi, Song, Jun, Yu, Cheng, Zheng, Bo, Huang, Gao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912973422854144
author Ni, Zanlin
Wang, Shenzhi
Yue, Yang
Yu, Tianyu
Zhao, Weilin
Hua, Yeguo
Chen, Tianyi
Song, Jun
Yu, Cheng
Zheng, Bo
Huang, Gao
author_facet Ni, Zanlin
Wang, Shenzhi
Yue, Yang
Yu, Tianyu
Zhao, Weilin
Hua, Yeguo
Chen, Tianyi
Song, Jun
Yu, Cheng
Zheng, Bo
Huang, Gao
contents Diffusion Large Language Models (dLLMs) break the rigid left-to-right constraint of traditional LLMs, enabling token generation in arbitrary orders. Intuitively, this flexibility implies a solution space that strictly supersets the fixed autoregressive trajectory, theoretically unlocking superior reasoning potential for general tasks like mathematics and coding. Consequently, numerous works have leveraged reinforcement learning (RL) to elicit the reasoning capability of dLLMs. In this paper, we reveal a counter-intuitive reality: arbitrary order generation, in its current form, narrows rather than expands the reasoning boundary of dLLMs. We find that dLLMs tend to exploit this order flexibility to bypass high-uncertainty tokens that are crucial for exploration, leading to a premature collapse of the solution space. This observation motivates a rethink of RL approaches for dLLMs, where considerable complexities, such as handling combinatorial trajectories and intractable likelihoods, are often devoted to preserving this flexibility. We demonstrate that effective reasoning can be better elicited by intentionally forgoing arbitrary order and applying standard Group Relative Policy Optimization (GRPO) instead. Our approach, JustGRPO, is minimalist yet surprisingly effective (e.g., 89.1% accuracy on GSM8K) while fully retaining the parallel decoding ability of dLLMs. Project page: https://nzl-thu.github.io/the-flexibility-trap
format Preprint
id arxiv_https___arxiv_org_abs_2601_15165
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle The Flexibility Trap: Why Arbitrary Order Limits Reasoning Potential in Diffusion Language Models
Ni, Zanlin
Wang, Shenzhi
Yue, Yang
Yu, Tianyu
Zhao, Weilin
Hua, Yeguo
Chen, Tianyi
Song, Jun
Yu, Cheng
Zheng, Bo
Huang, Gao
Computation and Language
Artificial Intelligence
Machine Learning
Diffusion Large Language Models (dLLMs) break the rigid left-to-right constraint of traditional LLMs, enabling token generation in arbitrary orders. Intuitively, this flexibility implies a solution space that strictly supersets the fixed autoregressive trajectory, theoretically unlocking superior reasoning potential for general tasks like mathematics and coding. Consequently, numerous works have leveraged reinforcement learning (RL) to elicit the reasoning capability of dLLMs. In this paper, we reveal a counter-intuitive reality: arbitrary order generation, in its current form, narrows rather than expands the reasoning boundary of dLLMs. We find that dLLMs tend to exploit this order flexibility to bypass high-uncertainty tokens that are crucial for exploration, leading to a premature collapse of the solution space. This observation motivates a rethink of RL approaches for dLLMs, where considerable complexities, such as handling combinatorial trajectories and intractable likelihoods, are often devoted to preserving this flexibility. We demonstrate that effective reasoning can be better elicited by intentionally forgoing arbitrary order and applying standard Group Relative Policy Optimization (GRPO) instead. Our approach, JustGRPO, is minimalist yet surprisingly effective (e.g., 89.1% accuracy on GSM8K) while fully retaining the parallel decoding ability of dLLMs. Project page: https://nzl-thu.github.io/the-flexibility-trap
title The Flexibility Trap: Why Arbitrary Order Limits Reasoning Potential in Diffusion Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2601.15165