TapOut: A Bandit-Based Approach to Dynamic Speculative Decoding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sridhar, Aditya, Sinnadurai, Nish, Lie, Sean, Thangarasa, Vithursan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908626172510208
author Sridhar, Aditya
Sinnadurai, Nish
Lie, Sean
Thangarasa, Vithursan
author_facet Sridhar, Aditya
Sinnadurai, Nish
Lie, Sean
Thangarasa, Vithursan
contents Speculative decoding accelerates LLMs by using a lightweight draft model to generate tokens autoregressively before verifying them in parallel with a larger target model. However, determining the optimal number of tokens to draft remains a key challenge limiting the approach's effectiveness. Dynamic speculative decoding aims to intelligently decide how many tokens to draft to achieve maximum speedups. Existing methods often rely on hand-tuned, sensitive thresholds (e.g., token entropy), which are costly to set and generalize poorly across models and domains. We propose TapOut, an online, training-free, plug-and-play algorithm for dynamic speculation policy selection using multi-armed bandits. Our approach employs a meta-algorithm that selects among multiple parameter-free dynamic speculation strategies based on past reward and exploration. We conduct extensive experiments across diverse model pairs and datasets, showing that TapOut achieves competitive or superior speedups compared to well-established dynamic speculation baselines without any hyperparameter tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2511_02017
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TapOut: A Bandit-Based Approach to Dynamic Speculative Decoding
Sridhar, Aditya
Sinnadurai, Nish
Lie, Sean
Thangarasa, Vithursan
Machine Learning
Computation and Language
Speculative decoding accelerates LLMs by using a lightweight draft model to generate tokens autoregressively before verifying them in parallel with a larger target model. However, determining the optimal number of tokens to draft remains a key challenge limiting the approach's effectiveness. Dynamic speculative decoding aims to intelligently decide how many tokens to draft to achieve maximum speedups. Existing methods often rely on hand-tuned, sensitive thresholds (e.g., token entropy), which are costly to set and generalize poorly across models and domains. We propose TapOut, an online, training-free, plug-and-play algorithm for dynamic speculation policy selection using multi-armed bandits. Our approach employs a meta-algorithm that selects among multiple parameter-free dynamic speculation strategies based on past reward and exploration. We conduct extensive experiments across diverse model pairs and datasets, showing that TapOut achieves competitive or superior speedups compared to well-established dynamic speculation baselines without any hyperparameter tuning.
title TapOut: A Bandit-Based Approach to Dynamic Speculative Decoding
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2511.02017