Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Huang, Jianuo, Zhang, Yaojie, Zhang, Qituan, Lin, Hao, Xu, Hanlin, Zhang, Linfeng
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911727239561216
author Huang, Jianuo
Zhang, Yaojie
Zhang, Qituan
Lin, Hao
Xu, Hanlin
Zhang, Linfeng
author_facet Huang, Jianuo
Zhang, Yaojie
Zhang, Qituan
Lin, Hao
Xu, Hanlin
Zhang, Linfeng
contents Speculative decoding accelerates LLM inference by drafting multiple tokens and verifying them in parallel with the target model. However, its practical speedup is constrained by the trade-off between draft quality and drafting cost: autoregressive drafters model causal dependencies among draft tokens but incur sequential overhead, while parallel drafters reduce drafting cost but weaken intra-block dependency modeling. In this paper, we propose Domino, a speculative decoding framework that decouples causal dependency modeling from expensive autoregressive draft execution. Domino first uses a parallel draft backbone to produce preliminary draft distributions for the entire block, and then applies a lightweight Domino head to refine them with prefix-dependent causal information. To stabilize teacher-forced causal encoding, we further introduce a base-anchored training curriculum that first strengthens the parallel backbone and then gradually shifts optimization toward the causally corrected final distribution. Experiments on Qwen3 models show that Domino achieves up to \(5.49\times\) end-to-end speedup under the Transformers backend and up to \(5.8\times\) throughput speedup under SGLang serving.
format Preprint
id arxiv_https___arxiv_org_abs_2605_29707
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding
Huang, Jianuo
Zhang, Yaojie
Zhang, Qituan
Lin, Hao
Xu, Hanlin
Zhang, Linfeng
Computation and Language
Speculative decoding accelerates LLM inference by drafting multiple tokens and verifying them in parallel with the target model. However, its practical speedup is constrained by the trade-off between draft quality and drafting cost: autoregressive drafters model causal dependencies among draft tokens but incur sequential overhead, while parallel drafters reduce drafting cost but weaken intra-block dependency modeling. In this paper, we propose Domino, a speculative decoding framework that decouples causal dependency modeling from expensive autoregressive draft execution. Domino first uses a parallel draft backbone to produce preliminary draft distributions for the entire block, and then applies a lightweight Domino head to refine them with prefix-dependent causal information. To stabilize teacher-forced causal encoding, we further introduce a base-anchored training curriculum that first strengthens the parallel backbone and then gradually shifts optimization toward the causally corrected final distribution. Experiments on Qwen3 models show that Domino achieves up to \(5.49\times\) end-to-end speedup under the Transformers backend and up to \(5.8\times\) throughput speedup under SGLang serving.
title Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding
topic Computation and Language
url https://arxiv.org/abs/2605.29707