RACER: Retrieval-Augmented Contextual Rapid Speculative Decoding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zihong, Li, Zuchao, Zhang, Lefei, Wang, Ping, Zhao, Hai
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914479636217856
author Zhang, Zihong
Li, Zuchao
Zhang, Lefei
Wang, Ping
Zhao, Hai
author_facet Zhang, Zihong
Li, Zuchao
Zhang, Lefei
Wang, Ping
Zhao, Hai
contents Autoregressive decoding in Large Language Models (LLMs) generates one token per step, causing high inference latency. Speculative decoding (SD) mitigates this through a guess-and-verify strategy, but existing training-free variants face trade-offs: retrieval-based drafts break when no exact match exists, while logits-based drafts lack structural guidance. We propose $\textbf{RACER}$ ($\textbf{R}$etrieval-$\textbf{A}$ugmented $\textbf{C}$ont$\textbf{e}$xtual $\textbf{R}$apid Speculative Decoding), a lightweight and training-free method that integrates retrieved exact patterns with logit-driven future cues. This unification supplies both reliable anchors and flexible extrapolation, yielding richer speculative drafts. Experiments on Spec-Bench, HumanEval, and MGSM-ZH demonstrate that RACER consistently accelerates inference, achieving more than $2\times$ speedup over autoregressive decoding, and outperforms prior training-free methods, offering a scalable, plug-and-play solution for efficient LLM decoding. Our source code is available at $\href{https://github.com/hkr04/RACER}{https://github.com/hkr04/RACER}$.
format Preprint
id arxiv_https___arxiv_org_abs_2604_14885
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RACER: Retrieval-Augmented Contextual Rapid Speculative Decoding
Zhang, Zihong
Li, Zuchao
Zhang, Lefei
Wang, Ping
Zhao, Hai
Computation and Language
Artificial Intelligence
Autoregressive decoding in Large Language Models (LLMs) generates one token per step, causing high inference latency. Speculative decoding (SD) mitigates this through a guess-and-verify strategy, but existing training-free variants face trade-offs: retrieval-based drafts break when no exact match exists, while logits-based drafts lack structural guidance. We propose $\textbf{RACER}$ ($\textbf{R}$etrieval-$\textbf{A}$ugmented $\textbf{C}$ont$\textbf{e}$xtual $\textbf{R}$apid Speculative Decoding), a lightweight and training-free method that integrates retrieved exact patterns with logit-driven future cues. This unification supplies both reliable anchors and flexible extrapolation, yielding richer speculative drafts. Experiments on Spec-Bench, HumanEval, and MGSM-ZH demonstrate that RACER consistently accelerates inference, achieving more than $2\times$ speedup over autoregressive decoding, and outperforms prior training-free methods, offering a scalable, plug-and-play solution for efficient LLM decoding. Our source code is available at $\href{https://github.com/hkr04/RACER}{https://github.com/hkr04/RACER}$.
title RACER: Retrieval-Augmented Contextual Rapid Speculative Decoding
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2604.14885