Localizing Paragraph Memorization in Language Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Stoehr, Niklas, Gordon, Mitchell, Zhang, Chiyuan, Lewis, Owen
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910389835399168
author Stoehr, Niklas
Gordon, Mitchell
Zhang, Chiyuan
Lewis, Owen
author_facet Stoehr, Niklas
Gordon, Mitchell
Zhang, Chiyuan
Lewis, Owen
contents Can we localize the weights and mechanisms used by a language model to memorize and recite entire paragraphs of its training data? In this paper, we show that while memorization is spread across multiple layers and model components, gradients of memorized paragraphs have a distinguishable spatial pattern, being larger in lower model layers than gradients of non-memorized examples. Moreover, the memorized examples can be unlearned by fine-tuning only the high-gradient weights. We localize a low-layer attention head that appears to be especially involved in paragraph memorization. This head is predominantly focusing its attention on distinctive, rare tokens that are least frequent in a corpus-level unigram distribution. Next, we study how localized memorization is across the tokens in the prefix by perturbing tokens and measuring the caused change in the decoding. A few distinctive tokens early in a prefix can often corrupt the entire continuation. Overall, memorized continuations are not only harder to unlearn, but also to corrupt than non-memorized ones.
format Preprint
id arxiv_https___arxiv_org_abs_2403_19851
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Localizing Paragraph Memorization in Language Models
Stoehr, Niklas
Gordon, Mitchell
Zhang, Chiyuan
Lewis, Owen
Computation and Language
Cryptography and Security
Machine Learning
Can we localize the weights and mechanisms used by a language model to memorize and recite entire paragraphs of its training data? In this paper, we show that while memorization is spread across multiple layers and model components, gradients of memorized paragraphs have a distinguishable spatial pattern, being larger in lower model layers than gradients of non-memorized examples. Moreover, the memorized examples can be unlearned by fine-tuning only the high-gradient weights. We localize a low-layer attention head that appears to be especially involved in paragraph memorization. This head is predominantly focusing its attention on distinctive, rare tokens that are least frequent in a corpus-level unigram distribution. Next, we study how localized memorization is across the tokens in the prefix by perturbing tokens and measuring the caused change in the decoding. A few distinctive tokens early in a prefix can often corrupt the entire continuation. Overall, memorized continuations are not only harder to unlearn, but also to corrupt than non-memorized ones.
title Localizing Paragraph Memorization in Language Models
topic Computation and Language
Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2403.19851