Guardado en:
Detalles Bibliográficos
Autores principales: Fang, Junjie, Tang, Likai, Bi, Hongzhe, Qin, Yujia, Sun, Si, Li, Zhenyu, Li, Haolun, Li, Yongjian, Cong, Xin, Lin, Yankai, Yan, Yukun, Shi, Xiaodong, Song, Sen, Liu, Zhiyuan, Sun, Maosong
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:https://arxiv.org/abs/2402.03009
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917752333139968
author Fang, Junjie
Tang, Likai
Bi, Hongzhe
Qin, Yujia
Sun, Si
Li, Zhenyu
Li, Haolun
Li, Yongjian
Cong, Xin
Lin, Yankai
Yan, Yukun
Shi, Xiaodong
Song, Sen
Liu, Zhiyuan
Sun, Maosong
author_facet Fang, Junjie
Tang, Likai
Bi, Hongzhe
Qin, Yujia
Sun, Si
Li, Zhenyu
Li, Haolun
Li, Yongjian
Cong, Xin
Lin, Yankai
Yan, Yukun
Shi, Xiaodong
Song, Sen
Liu, Zhiyuan
Sun, Maosong
contents Long-context processing is a critical ability that constrains the applicability of large language models (LLMs). Although there exist various methods devoted to enhancing the long-context processing ability of LLMs, they are developed in an isolated manner and lack systematic analysis and integration of their strengths, hindering further developments. In this paper, we introduce UniMem, a Unified framework that reformulates existing long-context methods from the view of Memory augmentation of LLMs. Distinguished by its four core dimensions-Memory Management, Memory Writing, Memory Reading, and Memory Injection, UniMem empowers researchers to conduct systematic exploration of long-context methods. We re-formulate 16 existing methods based on UniMem and analyze four representative methods: Transformer-XL, Memorizing Transformer, RMT, and Longformer into equivalent UniMem forms to reveal their design principles and strengths. Based on these analyses, we propose UniMix, an innovative approach that integrates the strengths of these algorithms. Experimental results show that UniMix achieves superior performance in handling long contexts with significantly lower perplexity than baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2402_03009
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle UniMem: Towards a Unified View of Long-Context Large Language Models
Fang, Junjie
Tang, Likai
Bi, Hongzhe
Qin, Yujia
Sun, Si
Li, Zhenyu
Li, Haolun
Li, Yongjian
Cong, Xin
Lin, Yankai
Yan, Yukun
Shi, Xiaodong
Song, Sen
Liu, Zhiyuan
Sun, Maosong
Computation and Language
Artificial Intelligence
Long-context processing is a critical ability that constrains the applicability of large language models (LLMs). Although there exist various methods devoted to enhancing the long-context processing ability of LLMs, they are developed in an isolated manner and lack systematic analysis and integration of their strengths, hindering further developments. In this paper, we introduce UniMem, a Unified framework that reformulates existing long-context methods from the view of Memory augmentation of LLMs. Distinguished by its four core dimensions-Memory Management, Memory Writing, Memory Reading, and Memory Injection, UniMem empowers researchers to conduct systematic exploration of long-context methods. We re-formulate 16 existing methods based on UniMem and analyze four representative methods: Transformer-XL, Memorizing Transformer, RMT, and Longformer into equivalent UniMem forms to reveal their design principles and strengths. Based on these analyses, we propose UniMix, an innovative approach that integrates the strengths of these algorithms. Experimental results show that UniMix achieves superior performance in handling long contexts with significantly lower perplexity than baselines.
title UniMem: Towards a Unified View of Long-Context Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2402.03009