Non-Markovian Discrete Diffusion with Causal Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yangtian, He, Sizhuang, Levine, Daniel, Zhao, Lawrence, Zhang, David, Rizvi, Syed A, Zhang, Shiyang, Zappala, Emanuele, Ying, Rex, van Dijk, David
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917048550948864
author Zhang, Yangtian
He, Sizhuang
Levine, Daniel
Zhao, Lawrence
Zhang, David
Rizvi, Syed A
Zhang, Shiyang
Zappala, Emanuele
Ying, Rex
van Dijk, David
author_facet Zhang, Yangtian
He, Sizhuang
Levine, Daniel
Zhao, Lawrence
Zhang, David
Rizvi, Syed A
Zhang, Shiyang
Zappala, Emanuele
Ying, Rex
van Dijk, David
contents Discrete diffusion models offer a flexible, controllable approach to structured sequence generation, yet they still lag behind causal language models in expressive power. A key limitation lies in their reliance on the Markovian assumption, which restricts each step to condition only on the current state, leading to potential uncorrectable error accumulation. In this paper, we introduce CaDDi (Causal Discrete Diffusion Model), a discrete diffusion model that conditions on the entire generative trajectory, thereby lifting the Markov constraint and allowing the model to revisit and improve past states. By unifying sequential (causal) and temporal (diffusion) reasoning in a single non-Markovian transformer, CaDDi also treats standard causal language models as a special case and permits the direct reuse of pretrained LLM weights with no architectural changes. Empirically, CaDDi outperforms state-of-the-art discrete diffusion baselines on natural-language benchmarks, substantially narrowing the remaining gap to large autoregressive transformers.
format Preprint
id arxiv_https___arxiv_org_abs_2502_09767
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Non-Markovian Discrete Diffusion with Causal Language Models
Zhang, Yangtian
He, Sizhuang
Levine, Daniel
Zhao, Lawrence
Zhang, David
Rizvi, Syed A
Zhang, Shiyang
Zappala, Emanuele
Ying, Rex
van Dijk, David
Machine Learning
Artificial Intelligence
Computation and Language
Discrete diffusion models offer a flexible, controllable approach to structured sequence generation, yet they still lag behind causal language models in expressive power. A key limitation lies in their reliance on the Markovian assumption, which restricts each step to condition only on the current state, leading to potential uncorrectable error accumulation. In this paper, we introduce CaDDi (Causal Discrete Diffusion Model), a discrete diffusion model that conditions on the entire generative trajectory, thereby lifting the Markov constraint and allowing the model to revisit and improve past states. By unifying sequential (causal) and temporal (diffusion) reasoning in a single non-Markovian transformer, CaDDi also treats standard causal language models as a special case and permits the direct reuse of pretrained LLM weights with no architectural changes. Empirically, CaDDi outperforms state-of-the-art discrete diffusion baselines on natural-language benchmarks, substantially narrowing the remaining gap to large autoregressive transformers.
title Non-Markovian Discrete Diffusion with Causal Language Models
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2502.09767