MetaState: Persistent Working Memory Enhances Reasoning in Discrete Diffusion Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xia, Kejing, Li, Mingzhe, Wei, Lixuan, Du, Zhenbang, Yuan, Xiangchi, Shi, Dachuan, Jin, Qirui, Lee, Wenke
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917368150622208
author Xia, Kejing
Li, Mingzhe
Wei, Lixuan
Du, Zhenbang
Yuan, Xiangchi
Shi, Dachuan
Jin, Qirui
Lee, Wenke
author_facet Xia, Kejing
Li, Mingzhe
Wei, Lixuan
Du, Zhenbang
Yuan, Xiangchi
Shi, Dachuan
Jin, Qirui
Lee, Wenke
contents Discrete diffusion language models (dLLMs) generate text by iteratively denoising a masked sequence. However, standard dLLMs condition each denoising step solely on the current hard-masked sequence, while intermediate continuous representations are discarded after sampling and remasking. We term this bottleneck the \textbf{Information Island} issue: continuous information remains isolated within individual denoising steps and fails to propagate across the trajectory. This bottleneck is especially harmful for reasoning, which requires intermediate reasoning state to be preserved and updated across many denoising steps. To address this limitation, we introduce \textbf{MetaState}, a lightweight recurrent augmentation that equips a frozen dLLM backbone with persistent, fixed-size working memory. MetaState comprises three modules with a shared time conditioner: a cross-attention \textbf{Mixer} that reads backbone activations into memory slots, a GRU-style \textbf{Updater} that integrates information across steps, and a cross-attention \textbf{Injector} that writes the updated memory back into the backbone. We train these modules with a dedicated $K$-step unrolling pipeline to learn multi-step dynamics. MetaState adds only ${\sim}0.6\%$ trainable parameters while keeping the backbone frozen, and consistently improves reasoning performance over frozen baselines on mathematical reasoning and code generation benchmarks, with an average gain of $4.5\%$ across all evaluations.
format Preprint
id arxiv_https___arxiv_org_abs_2603_01331
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MetaState: Persistent Working Memory Enhances Reasoning in Discrete Diffusion Language Models
Xia, Kejing
Li, Mingzhe
Wei, Lixuan
Du, Zhenbang
Yuan, Xiangchi
Shi, Dachuan
Jin, Qirui
Lee, Wenke
Computation and Language
Artificial Intelligence
Machine Learning
Discrete diffusion language models (dLLMs) generate text by iteratively denoising a masked sequence. However, standard dLLMs condition each denoising step solely on the current hard-masked sequence, while intermediate continuous representations are discarded after sampling and remasking. We term this bottleneck the \textbf{Information Island} issue: continuous information remains isolated within individual denoising steps and fails to propagate across the trajectory. This bottleneck is especially harmful for reasoning, which requires intermediate reasoning state to be preserved and updated across many denoising steps. To address this limitation, we introduce \textbf{MetaState}, a lightweight recurrent augmentation that equips a frozen dLLM backbone with persistent, fixed-size working memory. MetaState comprises three modules with a shared time conditioner: a cross-attention \textbf{Mixer} that reads backbone activations into memory slots, a GRU-style \textbf{Updater} that integrates information across steps, and a cross-attention \textbf{Injector} that writes the updated memory back into the backbone. We train these modules with a dedicated $K$-step unrolling pipeline to learn multi-step dynamics. MetaState adds only ${\sim}0.6\%$ trainable parameters while keeping the backbone frozen, and consistently improves reasoning performance over frozen baselines on mathematical reasoning and code generation benchmarks, with an average gain of $4.5\%$ across all evaluations.
title MetaState: Persistent Working Memory Enhances Reasoning in Discrete Diffusion Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2603.01331