The N-Grammys: Accelerating Autoregressive Inference with Learning-Free Batched Speculation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Stewart, Lawrence, Trager, Matthew, Gonugondla, Sujan Kumar, Soatto, Stefano
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909379060563968
author Stewart, Lawrence
Trager, Matthew
Gonugondla, Sujan Kumar
Soatto, Stefano
author_facet Stewart, Lawrence
Trager, Matthew
Gonugondla, Sujan Kumar
Soatto, Stefano
contents Speculative decoding aims to speed up autoregressive generation of a language model by verifying in parallel the tokens generated by a smaller draft model.In this work, we explore the effectiveness of learning-free, negligible-cost draft strategies, namely $N$-grams obtained from the model weights and the context. While the predicted next token of the base model is rarely the top prediction of these simple strategies, we observe that it is often within their top-$k$ predictions for small $k$. Based on this, we show that combinations of simple strategies can achieve significant inference speedups over different tasks. The overall performance is comparable to more complex methods, yet does not require expensive preprocessing or modification of the base model, and allows for seamless `plug-and-play' integration into pipelines.
format Preprint
id arxiv_https___arxiv_org_abs_2411_03786
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The N-Grammys: Accelerating Autoregressive Inference with Learning-Free Batched Speculation
Stewart, Lawrence
Trager, Matthew
Gonugondla, Sujan Kumar
Soatto, Stefano
Machine Learning
Speculative decoding aims to speed up autoregressive generation of a language model by verifying in parallel the tokens generated by a smaller draft model.In this work, we explore the effectiveness of learning-free, negligible-cost draft strategies, namely $N$-grams obtained from the model weights and the context. While the predicted next token of the base model is rarely the top prediction of these simple strategies, we observe that it is often within their top-$k$ predictions for small $k$. Based on this, we show that combinations of simple strategies can achieve significant inference speedups over different tasks. The overall performance is comparable to more complex methods, yet does not require expensive preprocessing or modification of the base model, and allows for seamless `plug-and-play' integration into pipelines.
title The N-Grammys: Accelerating Autoregressive Inference with Learning-Free Batched Speculation
topic Machine Learning
url https://arxiv.org/abs/2411.03786