Uncovering Latent Memories: Assessing Data Leakage and Memorization Patterns in Frontier AI Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Duan, Sunny, Khona, Mikail, Iyer, Abhiram, Schaeffer, Rylan, Fiete, Ila R
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913445779079168
author Duan, Sunny
Khona, Mikail
Iyer, Abhiram
Schaeffer, Rylan
Fiete, Ila R
author_facet Duan, Sunny
Khona, Mikail
Iyer, Abhiram
Schaeffer, Rylan
Fiete, Ila R
contents Frontier AI systems are making transformative impacts across society, but such benefits are not without costs: models trained on web-scale datasets containing personal and private data raise profound concerns about data privacy and security. Language models are trained on extensive corpora including potentially sensitive or proprietary information, and the risk of data leakage - where the model response reveals pieces of such information - remains inadequately understood. Prior work has investigated what factors drive memorization and have identified that sequence complexity and the number of repetitions drive memorization. Here, we focus on the evolution of memorization over training. We begin by reproducing findings that the probability of memorizing a sequence scales logarithmically with the number of times it is present in the data. We next show that sequences which are apparently not memorized after the first encounter can be "uncovered" throughout the course of training even without subsequent encounters, a phenomenon we term "latent memorization". The presence of latent memorization presents a challenge for data privacy as memorized sequences may be hidden at the final checkpoint of the model but remain easily recoverable. To this end, we develop a diagnostic test relying on the cross entropy loss to uncover latent memorized sequences with high accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2406_14549
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Uncovering Latent Memories: Assessing Data Leakage and Memorization Patterns in Frontier AI Models
Duan, Sunny
Khona, Mikail
Iyer, Abhiram
Schaeffer, Rylan
Fiete, Ila R
Computer Vision and Pattern Recognition
Machine Learning
Neurons and Cognition
Frontier AI systems are making transformative impacts across society, but such benefits are not without costs: models trained on web-scale datasets containing personal and private data raise profound concerns about data privacy and security. Language models are trained on extensive corpora including potentially sensitive or proprietary information, and the risk of data leakage - where the model response reveals pieces of such information - remains inadequately understood. Prior work has investigated what factors drive memorization and have identified that sequence complexity and the number of repetitions drive memorization. Here, we focus on the evolution of memorization over training. We begin by reproducing findings that the probability of memorizing a sequence scales logarithmically with the number of times it is present in the data. We next show that sequences which are apparently not memorized after the first encounter can be "uncovered" throughout the course of training even without subsequent encounters, a phenomenon we term "latent memorization". The presence of latent memorization presents a challenge for data privacy as memorized sequences may be hidden at the final checkpoint of the model but remain easily recoverable. To this end, we develop a diagnostic test relying on the cross entropy loss to uncover latent memorized sequences with high accuracy.
title Uncovering Latent Memories: Assessing Data Leakage and Memorization Patterns in Frontier AI Models
topic Computer Vision and Pattern Recognition
Machine Learning
Neurons and Cognition
url https://arxiv.org/abs/2406.14549