A Comprehensive Survey of Accelerated Generation Techniques in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khoshnoodi, Mahsa, Jain, Vinija, Gao, Mingye, Srikanth, Malavika, Chadha, Aman
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913360963960832
author Khoshnoodi, Mahsa
Jain, Vinija
Gao, Mingye
Srikanth, Malavika
Chadha, Aman
author_facet Khoshnoodi, Mahsa
Jain, Vinija
Gao, Mingye
Srikanth, Malavika
Chadha, Aman
contents Despite the crucial importance of accelerating text generation in large language models (LLMs) for efficiently producing content, the sequential nature of this process often leads to high inference latency, posing challenges for real-time applications. Various techniques have been proposed and developed to address these challenges and improve efficiency. This paper presents a comprehensive survey of accelerated generation techniques in autoregressive language models, aiming to understand the state-of-the-art methods and their applications. We categorize these techniques into several key areas: speculative decoding, early exiting mechanisms, and non-autoregressive methods. We discuss each category's underlying principles, advantages, limitations, and recent advancements. Through this survey, we aim to offer insights into the current landscape of techniques in LLMs and provide guidance for future research directions in this critical area of natural language processing.
format Preprint
id arxiv_https___arxiv_org_abs_2405_13019
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Comprehensive Survey of Accelerated Generation Techniques in Large Language Models
Khoshnoodi, Mahsa
Jain, Vinija
Gao, Mingye
Srikanth, Malavika
Chadha, Aman
Computation and Language
Artificial Intelligence
Despite the crucial importance of accelerating text generation in large language models (LLMs) for efficiently producing content, the sequential nature of this process often leads to high inference latency, posing challenges for real-time applications. Various techniques have been proposed and developed to address these challenges and improve efficiency. This paper presents a comprehensive survey of accelerated generation techniques in autoregressive language models, aiming to understand the state-of-the-art methods and their applications. We categorize these techniques into several key areas: speculative decoding, early exiting mechanisms, and non-autoregressive methods. We discuss each category's underlying principles, advantages, limitations, and recent advancements. Through this survey, we aim to offer insights into the current landscape of techniques in LLMs and provide guidance for future research directions in this critical area of natural language processing.
title A Comprehensive Survey of Accelerated Generation Techniques in Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2405.13019