Beyond the Target: From Imitation to Collaboration in Speculative Decoding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Jinze, Xu, Yixing, Li, Guanchen, Xu, Jinfeng, Yang, Shuo, Zhang, Yang, Yin, Xuanwu, Li, Dong, Ngai, Edith C. H., Barsoum, Emad
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917528633081856
author Li, Jinze
Xu, Yixing
Li, Guanchen
Xu, Jinfeng
Yang, Shuo
Zhang, Yang
Yin, Xuanwu
Li, Dong
Ngai, Edith C. H.
Barsoum, Emad
author_facet Li, Jinze
Xu, Yixing
Li, Guanchen
Xu, Jinfeng
Yang, Shuo
Zhang, Yang
Yin, Xuanwu
Li, Dong
Ngai, Edith C. H.
Barsoum, Emad
contents Speculative decoding (SPD) accelerates large language model (LLM) inference by letting a smaller draft model propose multiple future tokens that are verified in parallel by a larger target model. The dominant SPD paradigm treats the target model as the sole reliable teacher, accepting a draft token only when it exactly matches the target prediction. This design implicitly assumes that the target is always the better choice at every position. In practice, this assumption does not hold. Although the draft is the weaker model overall, it is not uniformly inferior at the token level. In a meaningful fraction of cases where draft and target disagree, the draft's choice is the one that leads to the correct final answer. Inspired by this, we introduce \textbf{Collaborative Speculative Decoding (CoSpec)}, a generalization of SPD that no longer treats the target model as the sole token-level authority. CoSpec trains an arbitration policy via reinforcement learning to decide whether to accept tokens from the draft or target model, selectively accepting draft tokens at mismatches when doing so is likely to yield a correct final answer. Experimental results show that CoSpec maintains substantial speedups while surpassing target-only performance. By shifting the emphasis from imitation to collaboration, CoSpec suggests a new perspective on speculative decoding.
format Preprint
id arxiv_https___arxiv_org_abs_2605_24793
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Beyond the Target: From Imitation to Collaboration in Speculative Decoding
Li, Jinze
Xu, Yixing
Li, Guanchen
Xu, Jinfeng
Yang, Shuo
Zhang, Yang
Yin, Xuanwu
Li, Dong
Ngai, Edith C. H.
Barsoum, Emad
Computation and Language
Speculative decoding (SPD) accelerates large language model (LLM) inference by letting a smaller draft model propose multiple future tokens that are verified in parallel by a larger target model. The dominant SPD paradigm treats the target model as the sole reliable teacher, accepting a draft token only when it exactly matches the target prediction. This design implicitly assumes that the target is always the better choice at every position. In practice, this assumption does not hold. Although the draft is the weaker model overall, it is not uniformly inferior at the token level. In a meaningful fraction of cases where draft and target disagree, the draft's choice is the one that leads to the correct final answer. Inspired by this, we introduce \textbf{Collaborative Speculative Decoding (CoSpec)}, a generalization of SPD that no longer treats the target model as the sole token-level authority. CoSpec trains an arbitration policy via reinforcement learning to decide whether to accept tokens from the draft or target model, selectively accepting draft tokens at mismatches when doing so is likely to yield a correct final answer. Experimental results show that CoSpec maintains substantial speedups while surpassing target-only performance. By shifting the emphasis from imitation to collaboration, CoSpec suggests a new perspective on speculative decoding.
title Beyond the Target: From Imitation to Collaboration in Speculative Decoding
topic Computation and Language
url https://arxiv.org/abs/2605.24793