TAPS: Task Aware Proposal Distributions for Speculative Sampling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zbib, Mohamad, Bazzi, Mohamad, Mohanna, Ammar, Hammoud, Hasan Abed Al Kader, Ghanem, Bernard
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915896580112384
author Zbib, Mohamad
Bazzi, Mohamad
Mohanna, Ammar
Hammoud, Hasan Abed Al Kader
Ghanem, Bernard
author_facet Zbib, Mohamad
Bazzi, Mohamad
Mohanna, Ammar
Hammoud, Hasan Abed Al Kader
Ghanem, Bernard
contents Speculative decoding accelerates autoregressive generation by letting a lightweight draft model propose future tokens that a larger target model then verifies in parallel. In practice, however, draft models are usually trained on broad generic corpora, which leaves it unclear how much speculative decoding quality depends on the draft training distribution. We study this question with lightweight HASS and EAGLE-2 drafters trained on MathInstruct, ShareGPT, and mixed-data variants, evaluated on MT-Bench, GSM8K, MATH-500, and SVAMP. Measured by acceptance length, task-specific training yields clear specialization: MathInstruct-trained drafts are strongest on reasoning benchmarks, while ShareGPT-trained drafts are strongest on MT-Bench. Mixed-data training improves robustness, but larger mixtures do not dominate across decoding temperatures. We also study how to combine specialized drafters at inference time. Naive checkpoint averaging performs poorly, whereas confidence-based routing improves over single-domain drafts and merged-tree verification yields the highest acceptance length overall for both backbones. Finally, confidence is a more useful routing signal than entropy: rejected tokens tend to have higher entropy, but confidence produces much clearer benchmark-level routing decisions. These results show that speculative decoding quality depends not only on draft architecture, but also on the match between draft training data and downstream workload, and that specialized drafters are better combined at inference time than in weight space.
format Preprint
id arxiv_https___arxiv_org_abs_2603_27027
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TAPS: Task Aware Proposal Distributions for Speculative Sampling
Zbib, Mohamad
Bazzi, Mohamad
Mohanna, Ammar
Hammoud, Hasan Abed Al Kader
Ghanem, Bernard
Computation and Language
Artificial Intelligence
Speculative decoding accelerates autoregressive generation by letting a lightweight draft model propose future tokens that a larger target model then verifies in parallel. In practice, however, draft models are usually trained on broad generic corpora, which leaves it unclear how much speculative decoding quality depends on the draft training distribution. We study this question with lightweight HASS and EAGLE-2 drafters trained on MathInstruct, ShareGPT, and mixed-data variants, evaluated on MT-Bench, GSM8K, MATH-500, and SVAMP. Measured by acceptance length, task-specific training yields clear specialization: MathInstruct-trained drafts are strongest on reasoning benchmarks, while ShareGPT-trained drafts are strongest on MT-Bench. Mixed-data training improves robustness, but larger mixtures do not dominate across decoding temperatures. We also study how to combine specialized drafters at inference time. Naive checkpoint averaging performs poorly, whereas confidence-based routing improves over single-domain drafts and merged-tree verification yields the highest acceptance length overall for both backbones. Finally, confidence is a more useful routing signal than entropy: rejected tokens tend to have higher entropy, but confidence produces much clearer benchmark-level routing decisions. These results show that speculative decoding quality depends not only on draft architecture, but also on the match between draft training data and downstream workload, and that specialized drafters are better combined at inference time than in weight space.
title TAPS: Task Aware Proposal Distributions for Speculative Sampling
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2603.27027