Draft on the Fly: Adaptive Self-Speculative Decoding using Cosine Similarity

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Metel, Michael R., Lu, Peng, Chen, Boxing, Rezagholizadeh, Mehdi, Kobyzev, Ivan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910628250124288
author Metel, Michael R.
Lu, Peng
Chen, Boxing
Rezagholizadeh, Mehdi
Kobyzev, Ivan
author_facet Metel, Michael R.
Lu, Peng
Chen, Boxing
Rezagholizadeh, Mehdi
Kobyzev, Ivan
contents We present a simple on the fly method for faster inference of large language models. Unlike other (self-)speculative decoding techniques, our method does not require fine-tuning or black-box optimization to generate a fixed draft model, relying instead on simple rules to generate varying draft models adapted to the input context. We show empirically that our light-weight algorithm is competitive with the current SOTA for self-speculative decoding, while being a truly plug-and-play method.
format Preprint
id arxiv_https___arxiv_org_abs_2410_01028
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Draft on the Fly: Adaptive Self-Speculative Decoding using Cosine Similarity
Metel, Michael R.
Lu, Peng
Chen, Boxing
Rezagholizadeh, Mehdi
Kobyzev, Ivan
Computation and Language
We present a simple on the fly method for faster inference of large language models. Unlike other (self-)speculative decoding techniques, our method does not require fine-tuning or black-box optimization to generate a fixed draft model, relying instead on simple rules to generate varying draft models adapted to the input context. We show empirically that our light-weight algorithm is competitive with the current SOTA for self-speculative decoding, while being a truly plug-and-play method.
title Draft on the Fly: Adaptive Self-Speculative Decoding using Cosine Similarity
topic Computation and Language
url https://arxiv.org/abs/2410.01028