Beyond the Limits: A Survey of Techniques to Extend the Context Length in Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Xindi, Salmani, Mahsa, Omidi, Parsa, Ren, Xiangyu, Rezagholizadeh, Mehdi, Eshaghi, Armaghan
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914815181586432
author Wang, Xindi
Salmani, Mahsa
Omidi, Parsa
Ren, Xiangyu
Rezagholizadeh, Mehdi
Eshaghi, Armaghan
author_facet Wang, Xindi
Salmani, Mahsa
Omidi, Parsa
Ren, Xiangyu
Rezagholizadeh, Mehdi
Eshaghi, Armaghan
contents Recently, large language models (LLMs) have shown remarkable capabilities including understanding context, engaging in logical reasoning, and generating responses. However, this is achieved at the expense of stringent computational and memory requirements, hindering their ability to effectively support long input sequences. This survey provides an inclusive review of the recent techniques and methods devised to extend the sequence length in LLMs, thereby enhancing their capacity for long-context understanding. In particular, we review and categorize a wide range of techniques including architectural modifications, such as modified positional encoding and altered attention mechanisms, which are designed to enhance the processing of longer sequences while avoiding a proportional increase in computational requirements. The diverse methodologies investigated in this study can be leveraged across different phases of LLMs, i.e., training, fine-tuning and inference. This enables LLMs to efficiently process extended sequences. The limitations of the current methodologies is discussed in the last section along with the suggestions for future research directions, underscoring the importance of sequence length in the continued advancement of LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2402_02244
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Beyond the Limits: A Survey of Techniques to Extend the Context Length in Large Language Models
Wang, Xindi
Salmani, Mahsa
Omidi, Parsa
Ren, Xiangyu
Rezagholizadeh, Mehdi
Eshaghi, Armaghan
Computation and Language
Machine Learning
Recently, large language models (LLMs) have shown remarkable capabilities including understanding context, engaging in logical reasoning, and generating responses. However, this is achieved at the expense of stringent computational and memory requirements, hindering their ability to effectively support long input sequences. This survey provides an inclusive review of the recent techniques and methods devised to extend the sequence length in LLMs, thereby enhancing their capacity for long-context understanding. In particular, we review and categorize a wide range of techniques including architectural modifications, such as modified positional encoding and altered attention mechanisms, which are designed to enhance the processing of longer sequences while avoiding a proportional increase in computational requirements. The diverse methodologies investigated in this study can be leveraged across different phases of LLMs, i.e., training, fine-tuning and inference. This enables LLMs to efficiently process extended sequences. The limitations of the current methodologies is discussed in the last section along with the suggestions for future research directions, underscoring the importance of sequence length in the continued advancement of LLMs.
title Beyond the Limits: A Survey of Techniques to Extend the Context Length in Large Language Models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2402.02244