Online Speculative Decoding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Xiaoxuan, Hu, Lanxiang, Bailis, Peter, Cheung, Alvin, Deng, Zhijie, Stoica, Ion, Zhang, Hao
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911911156645888
author Liu, Xiaoxuan
Hu, Lanxiang
Bailis, Peter
Cheung, Alvin
Deng, Zhijie
Stoica, Ion
Zhang, Hao
author_facet Liu, Xiaoxuan
Hu, Lanxiang
Bailis, Peter
Cheung, Alvin
Deng, Zhijie
Stoica, Ion
Zhang, Hao
contents Speculative decoding is a pivotal technique to accelerate the inference of large language models (LLMs) by employing a smaller draft model to predict the target model's outputs. However, its efficacy can be limited due to the low predictive accuracy of the draft model, particularly when faced with diverse text inputs and a significant capability gap between the draft and target models. We introduce online speculative decoding to address this challenge. The main idea is to continuously update the (multiple) draft model(s) on observed user query data. Adapting to query distribution mitigates the shifts between the training distribution of the draft model and the query distribution, enabling the draft model to more accurately predict the target model's outputs. We develop a prototype of online speculative decoding based on knowledge distillation and evaluate it using both synthetic and real query data. The results show a substantial increase in the token acceptance rate by 0.1 to 0.65, bringing 1.42x to 2.17x latency reduction. Our code is available at https://github.com/LiuXiaoxuanPKU/OSD.
format Preprint
id arxiv_https___arxiv_org_abs_2310_07177
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Online Speculative Decoding
Liu, Xiaoxuan
Hu, Lanxiang
Bailis, Peter
Cheung, Alvin
Deng, Zhijie
Stoica, Ion
Zhang, Hao
Artificial Intelligence
Computation and Language
Machine Learning
Speculative decoding is a pivotal technique to accelerate the inference of large language models (LLMs) by employing a smaller draft model to predict the target model's outputs. However, its efficacy can be limited due to the low predictive accuracy of the draft model, particularly when faced with diverse text inputs and a significant capability gap between the draft and target models. We introduce online speculative decoding to address this challenge. The main idea is to continuously update the (multiple) draft model(s) on observed user query data. Adapting to query distribution mitigates the shifts between the training distribution of the draft model and the query distribution, enabling the draft model to more accurately predict the target model's outputs. We develop a prototype of online speculative decoding based on knowledge distillation and evaluate it using both synthetic and real query data. The results show a substantial increase in the token acceptance rate by 0.1 to 0.65, bringing 1.42x to 2.17x latency reduction. Our code is available at https://github.com/LiuXiaoxuanPKU/OSD.
title Online Speculative Decoding
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2310.07177