Steering Pretrained Drafters during Speculative Decoding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Berdoz, Frédéric, Rheinboldt, Peer, Wattenhofer, Roger
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915614755389440
author Berdoz, Frédéric
Rheinboldt, Peer
Wattenhofer, Roger
author_facet Berdoz, Frédéric
Rheinboldt, Peer
Wattenhofer, Roger
contents Speculative decoding accelerates language model inference by separating generation into fast drafting and parallel verification. Its main limitation is drafter-verifier misalignment, which limits token acceptance and reduces overall effectiveness. While small drafting heads trained from scratch compensate with speed, they struggle when verification dominates latency or when inputs are out of distribution. In contrast, pretrained drafters, though slower, achieve higher acceptance rates thanks to stronger standalone generation capabilities, making them competitive when drafting latency is negligible relative to verification or communication overhead. In this work, we aim to improve the acceptance rates of pretrained drafters by introducing a lightweight dynamic alignment mechanism: a steering vector computed from the verifier's hidden states and injected into the pretrained drafter. Compared to existing offline alignment methods such as distillation, our approach boosts the number of accepted tokens by up to 35\% under standard sampling and 22\% under greedy sampling, all while incurring negligible computational overhead. Importantly, our approach can be retrofitted to existing architectures and pretrained models, enabling rapid adoption.
format Preprint
id arxiv_https___arxiv_org_abs_2511_09844
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Steering Pretrained Drafters during Speculative Decoding
Berdoz, Frédéric
Rheinboldt, Peer
Wattenhofer, Roger
Machine Learning
Performance
Speculative decoding accelerates language model inference by separating generation into fast drafting and parallel verification. Its main limitation is drafter-verifier misalignment, which limits token acceptance and reduces overall effectiveness. While small drafting heads trained from scratch compensate with speed, they struggle when verification dominates latency or when inputs are out of distribution. In contrast, pretrained drafters, though slower, achieve higher acceptance rates thanks to stronger standalone generation capabilities, making them competitive when drafting latency is negligible relative to verification or communication overhead. In this work, we aim to improve the acceptance rates of pretrained drafters by introducing a lightweight dynamic alignment mechanism: a steering vector computed from the verifier's hidden states and injected into the pretrained drafter. Compared to existing offline alignment methods such as distillation, our approach boosts the number of accepted tokens by up to 35\% under standard sampling and 22\% under greedy sampling, all while incurring negligible computational overhead. Importantly, our approach can be retrofitted to existing architectures and pretrained models, enabling rapid adoption.
title Steering Pretrained Drafters during Speculative Decoding
topic Machine Learning
Performance
url https://arxiv.org/abs/2511.09844