Retentive or Forgetful? Diving into the Knowledge Memorizing Mechanism of Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Cao, Boxi, Tang, Qiaoyu, Lin, Hongyu, Jiang, Shanshan, Dong, Bin, Han, Xianpei, Chen, Jiawei, Wang, Tianshu, Sun, Le
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913263690711040
author Cao, Boxi
Tang, Qiaoyu
Lin, Hongyu
Jiang, Shanshan
Dong, Bin
Han, Xianpei
Chen, Jiawei
Wang, Tianshu
Sun, Le
author_facet Cao, Boxi
Tang, Qiaoyu
Lin, Hongyu
Jiang, Shanshan
Dong, Bin
Han, Xianpei
Chen, Jiawei
Wang, Tianshu
Sun, Le
contents Memory is one of the most essential cognitive functions serving as a repository of world knowledge and episodes of activities. In recent years, large-scale pre-trained language models have shown remarkable memorizing ability. On the contrary, vanilla neural networks without pre-training have been long observed suffering from the catastrophic forgetting problem. To investigate such a retentive-forgetful contradiction and understand the memory mechanism of language models, we conduct thorough experiments by controlling the target knowledge types, the learning strategies and the learning schedules. We find that: 1) Vanilla language models are forgetful; 2) Pre-training leads to retentive language models; 3) Knowledge relevance and diversification significantly influence the memory formation. These conclusions are useful for understanding the abilities of pre-trained language models and shed light on designing and evaluating new learning and inference algorithms of language models.
format Preprint
id arxiv_https___arxiv_org_abs_2305_09144
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Retentive or Forgetful? Diving into the Knowledge Memorizing Mechanism of Language Models
Cao, Boxi
Tang, Qiaoyu
Lin, Hongyu
Jiang, Shanshan
Dong, Bin
Han, Xianpei
Chen, Jiawei
Wang, Tianshu
Sun, Le
Computation and Language
Artificial Intelligence
Memory is one of the most essential cognitive functions serving as a repository of world knowledge and episodes of activities. In recent years, large-scale pre-trained language models have shown remarkable memorizing ability. On the contrary, vanilla neural networks without pre-training have been long observed suffering from the catastrophic forgetting problem. To investigate such a retentive-forgetful contradiction and understand the memory mechanism of language models, we conduct thorough experiments by controlling the target knowledge types, the learning strategies and the learning schedules. We find that: 1) Vanilla language models are forgetful; 2) Pre-training leads to retentive language models; 3) Knowledge relevance and diversification significantly influence the memory formation. These conclusions are useful for understanding the abilities of pre-trained language models and shed light on designing and evaluating new learning and inference algorithms of language models.
title Retentive or Forgetful? Diving into the Knowledge Memorizing Mechanism of Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2305.09144