DReSD: Dense Retrieval for Speculative Decoding

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Gritta, Milan, Xue, Huiyin, Lampouras, Gerasimos
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915312303079424
author Gritta, Milan
Xue, Huiyin
Lampouras, Gerasimos
author_facet Gritta, Milan
Xue, Huiyin
Lampouras, Gerasimos
contents Speculative decoding (SD) accelerates Large Language Model (LLM) generation by using an efficient draft model to propose the next few tokens, which are verified by the LLM in a single forward call, reducing latency while preserving its outputs. We focus on retrieval-based SD where the draft model retrieves the next tokens from a non-parametric datastore. Sparse retrieval (REST), which operates on the surface form of strings, is currently the dominant paradigm due to its simplicity and scalability. However, its effectiveness is limited due to the usage of short contexts and exact string matching. Instead, we introduce Dense Retrieval for Speculative Decoding (DReSD), a novel framework that uses approximate nearest neighbour search with contextualised token embeddings to retrieve the most semantically relevant token sequences for SD. Extensive experiments show that DReSD achieves (on average) 87% higher acceptance rates, 65% longer accepted tokens and 19% faster generation speeds compared to sparse retrieval (REST).
format Preprint
id arxiv_https___arxiv_org_abs_2502_15572
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DReSD: Dense Retrieval for Speculative Decoding
Gritta, Milan
Xue, Huiyin
Lampouras, Gerasimos
Computation and Language
Speculative decoding (SD) accelerates Large Language Model (LLM) generation by using an efficient draft model to propose the next few tokens, which are verified by the LLM in a single forward call, reducing latency while preserving its outputs. We focus on retrieval-based SD where the draft model retrieves the next tokens from a non-parametric datastore. Sparse retrieval (REST), which operates on the surface form of strings, is currently the dominant paradigm due to its simplicity and scalability. However, its effectiveness is limited due to the usage of short contexts and exact string matching. Instead, we introduce Dense Retrieval for Speculative Decoding (DReSD), a novel framework that uses approximate nearest neighbour search with contextualised token embeddings to retrieve the most semantically relevant token sequences for SD. Extensive experiments show that DReSD achieves (on average) 87% higher acceptance rates, 65% longer accepted tokens and 19% faster generation speeds compared to sparse retrieval (REST).
title DReSD: Dense Retrieval for Speculative Decoding
topic Computation and Language
url https://arxiv.org/abs/2502.15572