EMS-SD: Efficient Multi-sample Speculative Decoding for Accelerating Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ni, Yunsheng, Liu, Chuanjian, Tang, Yehui, Han, Kai, Wang, Yunhe
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914971025145856
author Ni, Yunsheng
Liu, Chuanjian
Tang, Yehui
Han, Kai
Wang, Yunhe
author_facet Ni, Yunsheng
Liu, Chuanjian
Tang, Yehui
Han, Kai
Wang, Yunhe
contents Speculative decoding emerges as a pivotal technique for enhancing the inference speed of Large Language Models (LLMs). Despite recent research aiming to improve prediction efficiency, multi-sample speculative decoding has been overlooked due to varying numbers of accepted tokens within a batch in the verification phase. Vanilla method adds padding tokens in order to ensure that the number of new tokens remains consistent across samples. However, this increases the computational and memory access overhead, thereby reducing the speedup ratio. We propose a novel method that can resolve the issue of inconsistent tokens accepted by different samples without necessitating an increase in memory or computing overhead. Furthermore, our proposed method can handle the situation where the prediction tokens of different samples are inconsistent without the need to add padding tokens. Sufficient experiments demonstrate the efficacy of our method. Our code is available at https://github.com/niyunsheng/EMS-SD.
format Preprint
id arxiv_https___arxiv_org_abs_2405_07542
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EMS-SD: Efficient Multi-sample Speculative Decoding for Accelerating Large Language Models
Ni, Yunsheng
Liu, Chuanjian
Tang, Yehui
Han, Kai
Wang, Yunhe
Computation and Language
Speculative decoding emerges as a pivotal technique for enhancing the inference speed of Large Language Models (LLMs). Despite recent research aiming to improve prediction efficiency, multi-sample speculative decoding has been overlooked due to varying numbers of accepted tokens within a batch in the verification phase. Vanilla method adds padding tokens in order to ensure that the number of new tokens remains consistent across samples. However, this increases the computational and memory access overhead, thereby reducing the speedup ratio. We propose a novel method that can resolve the issue of inconsistent tokens accepted by different samples without necessitating an increase in memory or computing overhead. Furthermore, our proposed method can handle the situation where the prediction tokens of different samples are inconsistent without the need to add padding tokens. Sufficient experiments demonstrate the efficacy of our method. Our code is available at https://github.com/niyunsheng/EMS-SD.
title EMS-SD: Efficient Multi-sample Speculative Decoding for Accelerating Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2405.07542