Advancing Transformer Architecture in Long-Context Large Language Models: A Comprehensive Survey

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Huang, Yunpeng, Xu, Jingwei, Lai, Junyu, Jiang, Zixu, Chen, Taolue, Li, Zenan, Yao, Yuan, Ma, Xiaoxing, Yang, Lijuan, Chen, Hao, Li, Shupeng, Zhao, Penghao
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909119585189888
author Huang, Yunpeng
Xu, Jingwei
Lai, Junyu
Jiang, Zixu
Chen, Taolue
Li, Zenan
Yao, Yuan
Ma, Xiaoxing
Yang, Lijuan
Chen, Hao
Li, Shupeng
Zhao, Penghao
author_facet Huang, Yunpeng
Xu, Jingwei
Lai, Junyu
Jiang, Zixu
Chen, Taolue
Li, Zenan
Yao, Yuan
Ma, Xiaoxing
Yang, Lijuan
Chen, Hao
Li, Shupeng
Zhao, Penghao
contents Transformer-based Large Language Models (LLMs) have been applied in diverse areas such as knowledge bases, human interfaces, and dynamic agents, and marking a stride towards achieving Artificial General Intelligence (AGI). However, current LLMs are predominantly pretrained on short text snippets, which compromises their effectiveness in processing the long-context prompts that are frequently encountered in practical scenarios. This article offers a comprehensive survey of the recent advancement in Transformer-based LLM architectures aimed at enhancing the long-context capabilities of LLMs throughout the entire model lifecycle, from pre-training through to inference. We first delineate and analyze the problems of handling long-context input and output with the current Transformer-based models. We then provide a taxonomy and the landscape of upgrades on Transformer architecture to solve these problems. Afterwards, we provide an investigation on wildly used evaluation necessities tailored for long-context LLMs, including datasets, metrics, and baseline models, as well as optimization toolkits such as libraries, frameworks, and compilers to boost the efficacy of LLMs across different stages in runtime. Finally, we discuss the challenges and potential avenues for future research. A curated repository of relevant literature, continuously updated, is available at https://github.com/Strivin0311/long-llms-learning.
format Preprint
id arxiv_https___arxiv_org_abs_2311_12351
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Advancing Transformer Architecture in Long-Context Large Language Models: A Comprehensive Survey
Huang, Yunpeng
Xu, Jingwei
Lai, Junyu
Jiang, Zixu
Chen, Taolue
Li, Zenan
Yao, Yuan
Ma, Xiaoxing
Yang, Lijuan
Chen, Hao
Li, Shupeng
Zhao, Penghao
Computation and Language
Machine Learning
I.2.7; I.2.6; I.2.11
Transformer-based Large Language Models (LLMs) have been applied in diverse areas such as knowledge bases, human interfaces, and dynamic agents, and marking a stride towards achieving Artificial General Intelligence (AGI). However, current LLMs are predominantly pretrained on short text snippets, which compromises their effectiveness in processing the long-context prompts that are frequently encountered in practical scenarios. This article offers a comprehensive survey of the recent advancement in Transformer-based LLM architectures aimed at enhancing the long-context capabilities of LLMs throughout the entire model lifecycle, from pre-training through to inference. We first delineate and analyze the problems of handling long-context input and output with the current Transformer-based models. We then provide a taxonomy and the landscape of upgrades on Transformer architecture to solve these problems. Afterwards, we provide an investigation on wildly used evaluation necessities tailored for long-context LLMs, including datasets, metrics, and baseline models, as well as optimization toolkits such as libraries, frameworks, and compilers to boost the efficacy of LLMs across different stages in runtime. Finally, we discuss the challenges and potential avenues for future research. A curated repository of relevant literature, continuously updated, is available at https://github.com/Strivin0311/long-llms-learning.
title Advancing Transformer Architecture in Long-Context Large Language Models: A Comprehensive Survey
topic Computation and Language
Machine Learning
I.2.7; I.2.6; I.2.11
url https://arxiv.org/abs/2311.12351