$\text{Memory}^3$: Language Modeling with Explicit Memory

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yang, Hongkang, Lin, Zehao, Wang, Wenjin, Wu, Hao, Li, Zhiyu, Tang, Bo, Wei, Wenqiang, Wang, Jinbo, Tang, Zeyun, Song, Shichao, Xi, Chenyang, Yu, Yu, Chen, Kai, Xiong, Feiyu, Tang, Linpeng, E, Weinan
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909467882291200
author Yang, Hongkang
Lin, Zehao
Wang, Wenjin
Wu, Hao
Li, Zhiyu
Tang, Bo
Wei, Wenqiang
Wang, Jinbo
Tang, Zeyun
Song, Shichao
Xi, Chenyang
Yu, Yu
Chen, Kai
Xiong, Feiyu
Tang, Linpeng
E, Weinan
author_facet Yang, Hongkang
Lin, Zehao
Wang, Wenjin
Wu, Hao
Li, Zhiyu
Tang, Bo
Wei, Wenqiang
Wang, Jinbo
Tang, Zeyun
Song, Shichao
Xi, Chenyang
Yu, Yu
Chen, Kai
Xiong, Feiyu
Tang, Linpeng
E, Weinan
contents The training and inference of large language models (LLMs) are together a costly process that transports knowledge from raw data to meaningful computation. Inspired by the memory hierarchy of the human brain, we reduce this cost by equipping LLMs with explicit memory, a memory format cheaper than model parameters and text retrieval-augmented generation (RAG). Conceptually, with most of its knowledge externalized to explicit memories, the LLM can enjoy a smaller parameter size, training cost, and inference cost, all proportional to the amount of remaining "abstract knowledge". As a preliminary proof of concept, we train from scratch a 2.4B LLM, which achieves better performance than much larger LLMs as well as RAG models, and maintains higher decoding speed than RAG. The model is named $\text{Memory}^3$, since explicit memory is the third form of memory in LLMs after implicit memory (model parameters) and working memory (context key-values). We introduce a memory circuitry theory to support the externalization of knowledge, and present novel techniques including a memory sparsification mechanism that makes storage tractable and a two-stage pretraining scheme that facilitates memory formation.
format Preprint
id arxiv_https___arxiv_org_abs_2407_01178
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle $\text{Memory}^3$: Language Modeling with Explicit Memory
Yang, Hongkang
Lin, Zehao
Wang, Wenjin
Wu, Hao
Li, Zhiyu
Tang, Bo
Wei, Wenqiang
Wang, Jinbo
Tang, Zeyun
Song, Shichao
Xi, Chenyang
Yu, Yu
Chen, Kai
Xiong, Feiyu
Tang, Linpeng
E, Weinan
Computation and Language
Artificial Intelligence
Machine Learning
68T50
I.2.7
The training and inference of large language models (LLMs) are together a costly process that transports knowledge from raw data to meaningful computation. Inspired by the memory hierarchy of the human brain, we reduce this cost by equipping LLMs with explicit memory, a memory format cheaper than model parameters and text retrieval-augmented generation (RAG). Conceptually, with most of its knowledge externalized to explicit memories, the LLM can enjoy a smaller parameter size, training cost, and inference cost, all proportional to the amount of remaining "abstract knowledge". As a preliminary proof of concept, we train from scratch a 2.4B LLM, which achieves better performance than much larger LLMs as well as RAG models, and maintains higher decoding speed than RAG. The model is named $\text{Memory}^3$, since explicit memory is the third form of memory in LLMs after implicit memory (model parameters) and working memory (context key-values). We introduce a memory circuitry theory to support the externalization of knowledge, and present novel techniques including a memory sparsification mechanism that makes storage tractable and a two-stage pretraining scheme that facilitates memory formation.
title $\text{Memory}^3$: Language Modeling with Explicit Memory
topic Computation and Language
Artificial Intelligence
Machine Learning
68T50
I.2.7
url https://arxiv.org/abs/2407.01178