Gradient Localization Improves Lifelong Pretraining of Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fernandez, Jared, Bisk, Yonatan, Strubell, Emma
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929580782125056
author Fernandez, Jared
Bisk, Yonatan
Strubell, Emma
author_facet Fernandez, Jared
Bisk, Yonatan
Strubell, Emma
contents Large Language Models (LLMs) trained on web-scale text corpora have been shown to capture world knowledge in their parameters. However, the mechanism by which language models store different types of knowledge is poorly understood. In this work, we examine two types of knowledge relating to temporally sensitive entities and demonstrate that each type is localized to different sets of parameters within the LLMs. We hypothesize that the lack of consideration of the locality of knowledge in existing continual learning methods contributes to both: the failed uptake of new information, and catastrophic forgetting of previously learned information. We observe that sequences containing references to updated and newly mentioned entities exhibit larger gradient norms in a subset of layers. We demonstrate that targeting parameter updates to these relevant layers can improve the performance of continually pretraining on language containing temporal drift.
format Preprint
id arxiv_https___arxiv_org_abs_2411_04448
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Gradient Localization Improves Lifelong Pretraining of Language Models
Fernandez, Jared
Bisk, Yonatan
Strubell, Emma
Computation and Language
Large Language Models (LLMs) trained on web-scale text corpora have been shown to capture world knowledge in their parameters. However, the mechanism by which language models store different types of knowledge is poorly understood. In this work, we examine two types of knowledge relating to temporally sensitive entities and demonstrate that each type is localized to different sets of parameters within the LLMs. We hypothesize that the lack of consideration of the locality of knowledge in existing continual learning methods contributes to both: the failed uptake of new information, and catastrophic forgetting of previously learned information. We observe that sequences containing references to updated and newly mentioned entities exhibit larger gradient norms in a subset of layers. We demonstrate that targeting parameter updates to these relevant layers can improve the performance of continually pretraining on language containing temporal drift.
title Gradient Localization Improves Lifelong Pretraining of Language Models
topic Computation and Language
url https://arxiv.org/abs/2411.04448