Encoder-Decoder Diffusion Language Models for Efficient Training and Inference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Arriola, Marianne, Schiff, Yair, Phung, Hao, Gokaslan, Aaron, Kuleshov, Volodymyr
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914116233330688
author Arriola, Marianne
Schiff, Yair
Phung, Hao
Gokaslan, Aaron
Kuleshov, Volodymyr
author_facet Arriola, Marianne
Schiff, Yair
Phung, Hao
Gokaslan, Aaron
Kuleshov, Volodymyr
contents Discrete diffusion models enable parallel token sampling for faster inference than autoregressive approaches. However, prior diffusion models use a decoder-only architecture, which requires sampling algorithms that invoke the full network at every denoising step and incur high computational cost. Our key insight is that discrete diffusion models perform two types of computation: 1) representing clean tokens and 2) denoising corrupted tokens, which enables us to use separate modules for each task. We propose an encoder-decoder architecture to accelerate discrete diffusion inference, which relies on an encoder to represent clean tokens and a lightweight decoder to iteratively refine a noised sequence. We also show that this architecture enables faster training of block diffusion models, which partition sequences into blocks for better quality and are commonly used in diffusion language model inference. We introduce a framework for Efficient Encoder-Decoder Diffusion (E2D2), consisting of an architecture with specialized training and sampling algorithms, and we show that E2D2 achieves superior trade-offs between generation quality and inference throughput on summarization, translation, and mathematical reasoning tasks. We provide the code, model weights, and blog post on the project page: https://m-arriola.com/e2d2
format Preprint
id arxiv_https___arxiv_org_abs_2510_22852
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Encoder-Decoder Diffusion Language Models for Efficient Training and Inference
Arriola, Marianne
Schiff, Yair
Phung, Hao
Gokaslan, Aaron
Kuleshov, Volodymyr
Machine Learning
Artificial Intelligence
Discrete diffusion models enable parallel token sampling for faster inference than autoregressive approaches. However, prior diffusion models use a decoder-only architecture, which requires sampling algorithms that invoke the full network at every denoising step and incur high computational cost. Our key insight is that discrete diffusion models perform two types of computation: 1) representing clean tokens and 2) denoising corrupted tokens, which enables us to use separate modules for each task. We propose an encoder-decoder architecture to accelerate discrete diffusion inference, which relies on an encoder to represent clean tokens and a lightweight decoder to iteratively refine a noised sequence. We also show that this architecture enables faster training of block diffusion models, which partition sequences into blocks for better quality and are commonly used in diffusion language model inference. We introduce a framework for Efficient Encoder-Decoder Diffusion (E2D2), consisting of an architecture with specialized training and sampling algorithms, and we show that E2D2 achieves superior trade-offs between generation quality and inference throughput on summarization, translation, and mathematical reasoning tasks. We provide the code, model weights, and blog post on the project page: https://m-arriola.com/e2d2
title Encoder-Decoder Diffusion Language Models for Efficient Training and Inference
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2510.22852