Hierarchical Verification of Speculative Beams for Accelerating LLM Inference

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sen, Jaydip, Puvvala, Harshitha, Dasgupta, Subhasis
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910267721383936
author Sen, Jaydip
Puvvala, Harshitha
Dasgupta, Subhasis
author_facet Sen, Jaydip
Puvvala, Harshitha
Dasgupta, Subhasis
contents Large language models (LLMs) have achieved remarkable success across diverse natural language processing tasks but face persistent challenges in inference efficiency due to their autoregressive nature. While speculative decoding and beam sampling offer notable improvements, traditional methods verify draft sequences sequentially without prioritization, leading to unnecessary computational overhead. This work proposes the Hierarchical Verification Tree (HVT), a novel framework that restructures speculative beam decoding by prioritizing high-likelihood drafts and enabling early pruning of suboptimal candidates. Theoretical foundations and a formal verification-pruning algorithm are developed to ensure correctness and efficiency. Integration with standard LLM inference pipelines is achieved without requiring retraining or architecture modification. Experimental evaluations across multiple datasets and models demonstrate that HVT consistently outperforms existing speculative decoding schemes, achieving substantial reductions in inference time and energy consumption while maintaining or enhancing output quality. The findings highlight the potential of hierarchical verification strategies as a new direction for accelerating large language model inference.
format Preprint
id arxiv_https___arxiv_org_abs_2508_03726
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Hierarchical Verification of Speculative Beams for Accelerating LLM Inference
Sen, Jaydip
Puvvala, Harshitha
Dasgupta, Subhasis
Computation and Language
Large language models (LLMs) have achieved remarkable success across diverse natural language processing tasks but face persistent challenges in inference efficiency due to their autoregressive nature. While speculative decoding and beam sampling offer notable improvements, traditional methods verify draft sequences sequentially without prioritization, leading to unnecessary computational overhead. This work proposes the Hierarchical Verification Tree (HVT), a novel framework that restructures speculative beam decoding by prioritizing high-likelihood drafts and enabling early pruning of suboptimal candidates. Theoretical foundations and a formal verification-pruning algorithm are developed to ensure correctness and efficiency. Integration with standard LLM inference pipelines is achieved without requiring retraining or architecture modification. Experimental evaluations across multiple datasets and models demonstrate that HVT consistently outperforms existing speculative decoding schemes, achieving substantial reductions in inference time and energy consumption while maintaining or enhancing output quality. The findings highlight the potential of hierarchical verification strategies as a new direction for accelerating large language model inference.
title Hierarchical Verification of Speculative Beams for Accelerating LLM Inference
topic Computation and Language
url https://arxiv.org/abs/2508.03726