Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Teng, Yao, Wang, Fuyun, Liu, Xian, Chen, Zhekai, Shi, Han, Wang, Yu, Li, Zhenguo, Liu, Weiyang, Zou, Difan, Liu, Xihui
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909833884598272
author Teng, Yao
Wang, Fuyun
Liu, Xian
Chen, Zhekai
Shi, Han
Wang, Yu
Li, Zhenguo
Liu, Weiyang
Zou, Difan
Liu, Xihui
author_facet Teng, Yao
Wang, Fuyun
Liu, Xian
Chen, Zhekai
Shi, Han
Wang, Yu
Li, Zhenguo
Liu, Weiyang
Zou, Difan
Liu, Xihui
contents As a new paradigm of visual content generation, autoregressive text-to-image models suffer from slow inference due to their sequential token-by-token decoding process, often requiring thousands of model forward passes to generate a single image. To address this inefficiency, we propose Speculative Jacobi-Denoising Decoding (SJD2), a framework that incorporates the denoising process into Jacobi iterations to enable parallel token generation in autoregressive models. Our method introduces a next-clean-token prediction paradigm that enables the pre-trained autoregressive models to accept noise-perturbed token embeddings and predict the next clean tokens through low-cost fine-tuning. This denoising paradigm guides the model towards more stable Jacobi trajectories. During inference, our method initializes token sequences with Gaussian noise and performs iterative next-clean-token-prediction in the embedding space. We employ a probabilistic criterion to verify and accept multiple tokens in parallel, and refine the unaccepted tokens for the next iteration with the denoising trajectory. Experiments show that our method can accelerate generation by reducing model forward passes while maintaining the visual quality of generated images.
format Preprint
id arxiv_https___arxiv_org_abs_2510_08994
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image Generation
Teng, Yao
Wang, Fuyun
Liu, Xian
Chen, Zhekai
Shi, Han
Wang, Yu
Li, Zhenguo
Liu, Weiyang
Zou, Difan
Liu, Xihui
Computer Vision and Pattern Recognition
As a new paradigm of visual content generation, autoregressive text-to-image models suffer from slow inference due to their sequential token-by-token decoding process, often requiring thousands of model forward passes to generate a single image. To address this inefficiency, we propose Speculative Jacobi-Denoising Decoding (SJD2), a framework that incorporates the denoising process into Jacobi iterations to enable parallel token generation in autoregressive models. Our method introduces a next-clean-token prediction paradigm that enables the pre-trained autoregressive models to accept noise-perturbed token embeddings and predict the next clean tokens through low-cost fine-tuning. This denoising paradigm guides the model towards more stable Jacobi trajectories. During inference, our method initializes token sequences with Gaussian noise and performs iterative next-clean-token-prediction in the embedding space. We employ a probabilistic criterion to verify and accept multiple tokens in parallel, and refine the unaccepted tokens for the next iteration with the denoising trajectory. Experiments show that our method can accelerate generation by reducing model forward passes while maintaining the visual quality of generated images.
title Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.08994