Thinking Inside the Mask: In-Place Prompting in Diffusion LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jin, Xiangqi, Wang, Yuxuan, Gao, Yifeng, Wen, Zichen, Qi, Biqing, Liu, Dongrui, Zhang, Linfeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915545501138944
author Jin, Xiangqi
Wang, Yuxuan
Gao, Yifeng
Wen, Zichen
Qi, Biqing
Liu, Dongrui
Zhang, Linfeng
author_facet Jin, Xiangqi
Wang, Yuxuan
Gao, Yifeng
Wen, Zichen
Qi, Biqing
Liu, Dongrui
Zhang, Linfeng
contents Despite large language models (LLMs) have achieved remarkable success, their prefix-only prompting paradigm and sequential generation process offer limited flexibility for bidirectional information. Diffusion large language models (dLLMs) present new opportunities through their bidirectional attention mechanisms and iterative refinement processes, enabling more flexible in-place prompting strategies. We introduce ICE (In-Place Chain-of-Thought Prompting with Early Exit), a novel framework that transforms prefix-only prompting into in-place prompting specifically designed for dLLMs. ICE integrates in-place prompts directly within masked token positions during iterative refinement and employs a confidence-aware early exit mechanism to significantly reduce computational overhead. Extensive experiments demonstrate ICE's effectiveness, achieving up to 17.29% accuracy improvement with 4.12$\times$ speedup on GSM8K, and up to 276.67$\times$ acceleration on MMLU while maintaining competitive performance.
format Preprint
id arxiv_https___arxiv_org_abs_2508_10736
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Thinking Inside the Mask: In-Place Prompting in Diffusion LLMs
Jin, Xiangqi
Wang, Yuxuan
Gao, Yifeng
Wen, Zichen
Qi, Biqing
Liu, Dongrui
Zhang, Linfeng
Computation and Language
Despite large language models (LLMs) have achieved remarkable success, their prefix-only prompting paradigm and sequential generation process offer limited flexibility for bidirectional information. Diffusion large language models (dLLMs) present new opportunities through their bidirectional attention mechanisms and iterative refinement processes, enabling more flexible in-place prompting strategies. We introduce ICE (In-Place Chain-of-Thought Prompting with Early Exit), a novel framework that transforms prefix-only prompting into in-place prompting specifically designed for dLLMs. ICE integrates in-place prompts directly within masked token positions during iterative refinement and employs a confidence-aware early exit mechanism to significantly reduce computational overhead. Extensive experiments demonstrate ICE's effectiveness, achieving up to 17.29% accuracy improvement with 4.12$\times$ speedup on GSM8K, and up to 276.67$\times$ acceleration on MMLU while maintaining competitive performance.
title Thinking Inside the Mask: In-Place Prompting in Diffusion LLMs
topic Computation and Language
url https://arxiv.org/abs/2508.10736