Learning to Draft: Adaptive Speculative Decoding with Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Jiebin, Yu, Zhenghan, Wang, Liang, Yang, Nan, Yu, Eugene J., Li, Zheng, Song, Yifan, Zhu, Dawei, Zhang, Xingxing, Wei, Furu, Li, Sujian
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915827635191808
author Zhang, Jiebin
Yu, Zhenghan
Wang, Liang
Yang, Nan
Yu, Eugene J.
Li, Zheng
Song, Yifan
Zhu, Dawei
Zhang, Xingxing
Wei, Furu
Li, Sujian
author_facet Zhang, Jiebin
Yu, Zhenghan
Wang, Liang
Yang, Nan
Yu, Eugene J.
Li, Zheng
Song, Yifan
Zhu, Dawei
Zhang, Xingxing
Wei, Furu
Li, Sujian
contents Speculative decoding accelerates large language model (LLM) inference by using a small draft model to generate candidate tokens for a larger target model to verify. The efficacy of this technique hinges on the trade-off between the time spent on drafting candidates and verifying them. However, current state-of-the-art methods rely on a static time allocation, while recent dynamic approaches optimize for proxy metrics like acceptance length, often neglecting the true time cost and treating the drafting and verification phases in isolation. To address these limitations, we introduce Learning to Draft (LTD), a novel method that directly optimizes for throughput of each draft-and-verify cycle. We formulate the problem as a reinforcement learning environment and train two co-adaptive policies to dynamically coordinate the draft and verification phases. This encourages the policies to adapt to each other and explicitly maximize decoding efficiency. We conducted extensive evaluations on five diverse LLMs and four distinct tasks. Our results show that LTD achieves speedup ratios ranging from 2.24x to 4.32x, outperforming the state-of-the-art method Eagle3 up to 36.4%.
format Preprint
id arxiv_https___arxiv_org_abs_2603_01639
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learning to Draft: Adaptive Speculative Decoding with Reinforcement Learning
Zhang, Jiebin
Yu, Zhenghan
Wang, Liang
Yang, Nan
Yu, Eugene J.
Li, Zheng
Song, Yifan
Zhu, Dawei
Zhang, Xingxing
Wei, Furu
Li, Sujian
Computation and Language
Speculative decoding accelerates large language model (LLM) inference by using a small draft model to generate candidate tokens for a larger target model to verify. The efficacy of this technique hinges on the trade-off between the time spent on drafting candidates and verifying them. However, current state-of-the-art methods rely on a static time allocation, while recent dynamic approaches optimize for proxy metrics like acceptance length, often neglecting the true time cost and treating the drafting and verification phases in isolation. To address these limitations, we introduce Learning to Draft (LTD), a novel method that directly optimizes for throughput of each draft-and-verify cycle. We formulate the problem as a reinforcement learning environment and train two co-adaptive policies to dynamically coordinate the draft and verification phases. This encourages the policies to adapt to each other and explicitly maximize decoding efficiency. We conducted extensive evaluations on five diverse LLMs and four distinct tasks. Our results show that LTD achieves speedup ratios ranging from 2.24x to 4.32x, outperforming the state-of-the-art method Eagle3 up to 36.4%.
title Learning to Draft: Adaptive Speculative Decoding with Reinforcement Learning
topic Computation and Language
url https://arxiv.org/abs/2603.01639