Guardado en:
Detalles Bibliográficos
Autores principales: Williams, Miles, Kwon, Young D., Li, Rui, Kouris, Alexandros, Venieris, Stylianos I.
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:https://arxiv.org/abs/2602.13836
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912904515682304
author Williams, Miles
Kwon, Young D.
Li, Rui
Kouris, Alexandros
Venieris, Stylianos I.
author_facet Williams, Miles
Kwon, Young D.
Li, Rui
Kouris, Alexandros
Venieris, Stylianos I.
contents Speculative decoding has rapidly emerged as a leading approach for accelerating language model (LM) inference, as it offers substantial speedups while yielding identical outputs. This relies upon a small draft model, tasked with predicting the outputs of the target model. State-of-the-art speculative decoding methods use a draft model consisting of a single decoder layer and output embedding matrix, with the latter dominating drafting time for the latest LMs. Recent work has sought to address this output distribution bottleneck by reducing the vocabulary of the draft model. Although this can improve throughput, it compromises speculation effectiveness when the target token is out-of-vocabulary. In this paper, we argue for vocabulary speculation as an alternative to a reduced vocabulary. We propose SpecVocab, an efficient and effective method that selects a vocabulary subset per decoding step. Across a variety of tasks, we demonstrate that SpecVocab can achieve a higher acceptance length than state-of-the-art speculative decoding approach, EAGLE-3. Notably, this yields up to an 8.1% increase in average throughput over EAGLE-3.
format Preprint
id arxiv_https___arxiv_org_abs_2602_13836
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Speculative Decoding with a Speculative Vocabulary
Williams, Miles
Kwon, Young D.
Li, Rui
Kouris, Alexandros
Venieris, Stylianos I.
Computation and Language
Speculative decoding has rapidly emerged as a leading approach for accelerating language model (LM) inference, as it offers substantial speedups while yielding identical outputs. This relies upon a small draft model, tasked with predicting the outputs of the target model. State-of-the-art speculative decoding methods use a draft model consisting of a single decoder layer and output embedding matrix, with the latter dominating drafting time for the latest LMs. Recent work has sought to address this output distribution bottleneck by reducing the vocabulary of the draft model. Although this can improve throughput, it compromises speculation effectiveness when the target token is out-of-vocabulary. In this paper, we argue for vocabulary speculation as an alternative to a reduced vocabulary. We propose SpecVocab, an efficient and effective method that selects a vocabulary subset per decoding step. Across a variety of tasks, we demonstrate that SpecVocab can achieve a higher acceptance length than state-of-the-art speculative decoding approach, EAGLE-3. Notably, this yields up to an 8.1% increase in average throughput over EAGLE-3.
title Speculative Decoding with a Speculative Vocabulary
topic Computation and Language
url https://arxiv.org/abs/2602.13836