HADES: Hardware Accelerated Decoding for Efficient Speculation in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Ze, Jin, Yihong, Xu, Xinhe
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910781776330752
author Yang, Ze
Jin, Yihong
Xu, Xinhe
author_facet Yang, Ze
Jin, Yihong
Xu, Xinhe
contents Large Language Models (LLMs) have revolutionized natural language processing by understanding and generating human-like text. However, the increasing demand for more sophisticated LLMs presents significant computational challenges due to their scale and complexity. This paper introduces Hardware Accelerated Decoding (HADES), a novel approach to enhance the performance and energy efficiency of LLMs. We address the design of an LLM accelerator with hardware-level speculative decoding support, a concept not previously explored in existing literature. Our work demonstrates how speculative decoding can significantly improve the efficiency of LLM operations, paving the way for more advanced and practical applications of these models.
format Preprint
id arxiv_https___arxiv_org_abs_2412_19925
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle HADES: Hardware Accelerated Decoding for Efficient Speculation in Large Language Models
Yang, Ze
Jin, Yihong
Xu, Xinhe
Computation and Language
Artificial Intelligence
Hardware Architecture
Large Language Models (LLMs) have revolutionized natural language processing by understanding and generating human-like text. However, the increasing demand for more sophisticated LLMs presents significant computational challenges due to their scale and complexity. This paper introduces Hardware Accelerated Decoding (HADES), a novel approach to enhance the performance and energy efficiency of LLMs. We address the design of an LLM accelerator with hardware-level speculative decoding support, a concept not previously explored in existing literature. Our work demonstrates how speculative decoding can significantly improve the efficiency of LLM operations, paving the way for more advanced and practical applications of these models.
title HADES: Hardware Accelerated Decoding for Efficient Speculation in Large Language Models
topic Computation and Language
Artificial Intelligence
Hardware Architecture
url https://arxiv.org/abs/2412.19925