Accelerating Diffusion LLMs via Adaptive Parallel Decoding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Israel, Daniel, Broeck, Guy Van den, Grover, Aditya
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911241916645376
author Israel, Daniel
Broeck, Guy Van den
Grover, Aditya
author_facet Israel, Daniel
Broeck, Guy Van den
Grover, Aditya
contents The generation speed of LLMs are bottlenecked by autoregressive decoding, where tokens are predicted sequentially one by one. Alternatively, diffusion large language models (dLLMs) theoretically allow for parallel token generation, but in practice struggle to achieve the speed of autoregressive models without significantly sacrificing quality. We therefore introduce adaptive parallel decoding (APD), a novel method that dynamically adjusts the number of tokens sampled in parallel. We achieve this by defining a multiplicative mixture between the dLLM marginal probabilities and the joint probability of sequences under a small auxiliary autoregressive model. This inverts the standard setup of speculative decoding, where the goal is to sample from a large autoregressive verifier by drafting from a smaller model. We further optimize APD by enabling KV caching and limiting the size of the masked input. Altogether, our method puts forward three tunable parameters to flexibly tradeoff throughput and quality. We show that APD provides markedly higher throughput with minimal quality degradations on downstream benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00413
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Accelerating Diffusion LLMs via Adaptive Parallel Decoding
Israel, Daniel
Broeck, Guy Van den
Grover, Aditya
Computation and Language
Artificial Intelligence
Machine Learning
Performance
The generation speed of LLMs are bottlenecked by autoregressive decoding, where tokens are predicted sequentially one by one. Alternatively, diffusion large language models (dLLMs) theoretically allow for parallel token generation, but in practice struggle to achieve the speed of autoregressive models without significantly sacrificing quality. We therefore introduce adaptive parallel decoding (APD), a novel method that dynamically adjusts the number of tokens sampled in parallel. We achieve this by defining a multiplicative mixture between the dLLM marginal probabilities and the joint probability of sequences under a small auxiliary autoregressive model. This inverts the standard setup of speculative decoding, where the goal is to sample from a large autoregressive verifier by drafting from a smaller model. We further optimize APD by enabling KV caching and limiting the size of the masked input. Altogether, our method puts forward three tunable parameters to flexibly tradeoff throughput and quality. We show that APD provides markedly higher throughput with minimal quality degradations on downstream benchmarks.
title Accelerating Diffusion LLMs via Adaptive Parallel Decoding
topic Computation and Language
Artificial Intelligence
Machine Learning
Performance
url https://arxiv.org/abs/2506.00413