Speculative Decoding via Early-exiting for Faster LLM Inference with Thompson Sampling Control Mechanism

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Jiahao, Wang, Qifan, Wang, Jingang, Cai, Xunliang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911908013015040
author Liu, Jiahao
Wang, Qifan
Wang, Jingang
Cai, Xunliang
author_facet Liu, Jiahao
Wang, Qifan
Wang, Jingang
Cai, Xunliang
contents The recent advancements in large language models (LLMs) have been extraordinary, yet the escalating inference costs associated with them present challenges in real-world applications. To address these challenges, we propose a novel approach called Early-exiting Speculative Decoding (EESD) with lossless acceleration. Specifically, EESD utilizes a segment of the LLM to generate draft tokens, incorporating Early-exiting structures after the first N layers. To enhance the quality of draft tokens, a self-distillation method is integrated. This early-exiting design not only reduces deployment and training costs but also significantly accelerates the token generation speed. Moreover, we introduce a novel sampling mechanism that leverages Thompson Sampling to regulate the generation processes, automatically determining the quantity of draft tokens in each round. The original LLM is then employed to validate these draft tokens through a single forward pass, and thus guarantees that the final output text maintains a distribution consistent with vanilla auto-regressive decoding. The experimental results on both 13B and 70B models demonstrate that our approach decodes tokens at a markedly accelerated rate compared to prior methods, showing the effectiveness of our approach.
format Preprint
id arxiv_https___arxiv_org_abs_2406_03853
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Speculative Decoding via Early-exiting for Faster LLM Inference with Thompson Sampling Control Mechanism
Liu, Jiahao
Wang, Qifan
Wang, Jingang
Cai, Xunliang
Computation and Language
The recent advancements in large language models (LLMs) have been extraordinary, yet the escalating inference costs associated with them present challenges in real-world applications. To address these challenges, we propose a novel approach called Early-exiting Speculative Decoding (EESD) with lossless acceleration. Specifically, EESD utilizes a segment of the LLM to generate draft tokens, incorporating Early-exiting structures after the first N layers. To enhance the quality of draft tokens, a self-distillation method is integrated. This early-exiting design not only reduces deployment and training costs but also significantly accelerates the token generation speed. Moreover, we introduce a novel sampling mechanism that leverages Thompson Sampling to regulate the generation processes, automatically determining the quantity of draft tokens in each round. The original LLM is then employed to validate these draft tokens through a single forward pass, and thus guarantees that the final output text maintains a distribution consistent with vanilla auto-regressive decoding. The experimental results on both 13B and 70B models demonstrate that our approach decodes tokens at a markedly accelerated rate compared to prior methods, showing the effectiveness of our approach.
title Speculative Decoding via Early-exiting for Faster LLM Inference with Thompson Sampling Control Mechanism
topic Computation and Language
url https://arxiv.org/abs/2406.03853