SDAR: A Synergistic Diffusion-AutoRegression Paradigm for Scalable Sequence Generation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Cheng, Shuang, Bian, Yihan, Liu, Dawei, Zhang, Linfeng, Yao, Qian, Tian, Zhongbo, Wang, Wenhai, Guo, Qipeng, Chen, Kai, Qi, Biqing, Zhou, Bowen
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911219293618176
author Cheng, Shuang
Bian, Yihan
Liu, Dawei
Zhang, Linfeng
Yao, Qian
Tian, Zhongbo
Wang, Wenhai
Guo, Qipeng
Chen, Kai
Qi, Biqing
Zhou, Bowen
author_facet Cheng, Shuang
Bian, Yihan
Liu, Dawei
Zhang, Linfeng
Yao, Qian
Tian, Zhongbo
Wang, Wenhai
Guo, Qipeng
Chen, Kai
Qi, Biqing
Zhou, Bowen
contents We propose SDAR, a Synergistic Diffusion-Autoregression paradigm that unifies the training efficiency of autoregressive models with the parallel inference capability of diffusion. Instead of costly end-to-end diffusion training, SDAR performs a lightweight paradigm conversion that transforms a well-trained autoregressive (AR) model into a blockwise diffusion model through brief, data-efficient adaptation. During inference, SDAR generates sequences autoregressively across blocks for global coherence while decoding all tokens within each block in parallel via a discrete diffusion process. Extensive experiments show that AR models remain substantially more compute-efficient than masked diffusion models, providing a strong foundation for adaptation. Building on this insight, SDAR achieves efficient AR-to-diffusion conversion with minimal cost, preserving AR-level performance while enabling parallel generation. Scaling studies across dense and Mixture-of-Experts architectures confirm that SDAR scales without compromise: larger models exhibit stronger robustness to block size and decoding thresholds, yielding greater speedups without accuracy loss. Beyond efficiency, SDAR demonstrates enhanced reasoning and domain adaptability. Our 30B MoE model surpasses its AR counterpart on challenging scientific reasoning benchmarks such as GPQA and ChemBench, and gains further improvements under test-time scaling methods like majority voting and pass@k. Together, these results establish SDAR as a practical paradigm that combines the strengths of autoregression and diffusion for scalable, high-throughput reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2510_06303
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SDAR: A Synergistic Diffusion-AutoRegression Paradigm for Scalable Sequence Generation
Cheng, Shuang
Bian, Yihan
Liu, Dawei
Zhang, Linfeng
Yao, Qian
Tian, Zhongbo
Wang, Wenhai
Guo, Qipeng
Chen, Kai
Qi, Biqing
Zhou, Bowen
Machine Learning
Artificial Intelligence
We propose SDAR, a Synergistic Diffusion-Autoregression paradigm that unifies the training efficiency of autoregressive models with the parallel inference capability of diffusion. Instead of costly end-to-end diffusion training, SDAR performs a lightweight paradigm conversion that transforms a well-trained autoregressive (AR) model into a blockwise diffusion model through brief, data-efficient adaptation. During inference, SDAR generates sequences autoregressively across blocks for global coherence while decoding all tokens within each block in parallel via a discrete diffusion process. Extensive experiments show that AR models remain substantially more compute-efficient than masked diffusion models, providing a strong foundation for adaptation. Building on this insight, SDAR achieves efficient AR-to-diffusion conversion with minimal cost, preserving AR-level performance while enabling parallel generation. Scaling studies across dense and Mixture-of-Experts architectures confirm that SDAR scales without compromise: larger models exhibit stronger robustness to block size and decoding thresholds, yielding greater speedups without accuracy loss. Beyond efficiency, SDAR demonstrates enhanced reasoning and domain adaptability. Our 30B MoE model surpasses its AR counterpart on challenging scientific reasoning benchmarks such as GPQA and ChemBench, and gains further improvements under test-time scaling methods like majority voting and pass@k. Together, these results establish SDAR as a practical paradigm that combines the strengths of autoregression and diffusion for scalable, high-throughput reasoning.
title SDAR: A Synergistic Diffusion-AutoRegression Paradigm for Scalable Sequence Generation
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2510.06303