Saved in:
Bibliographic Details
Main Author: Arnold, Stefan
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2506.09591
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911000135991296
author Arnold, Stefan
author_facet Arnold, Stefan
contents Language Models (LMs) are prone to memorizing parts of their data during training and unintentionally emitting them at generation time, raising concerns about privacy leakage and disclosure of intellectual property. While previous research has identified properties such as context length, parameter size, and duplication frequency, as key drivers of unintended memorization, little is known about how the latent structure modulates this rate of memorization. We investigate the role of Intrinsic Dimension (ID), a geometric proxy for the structural complexity of a sequence in latent space, in modulating memorization. Our findings suggest that ID acts as a suppressive signal for memorization: compared to low-ID sequences, high-ID sequences are less likely to be memorized, particularly in overparameterized models and under sparse exposure. These findings highlight the interaction between scale, exposure, and complexity in shaping memorization.
format Preprint
id arxiv_https___arxiv_org_abs_2506_09591
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Memorization in Language Models through the Lens of Intrinsic Dimension
Arnold, Stefan
Computation and Language
Language Models (LMs) are prone to memorizing parts of their data during training and unintentionally emitting them at generation time, raising concerns about privacy leakage and disclosure of intellectual property. While previous research has identified properties such as context length, parameter size, and duplication frequency, as key drivers of unintended memorization, little is known about how the latent structure modulates this rate of memorization. We investigate the role of Intrinsic Dimension (ID), a geometric proxy for the structural complexity of a sequence in latent space, in modulating memorization. Our findings suggest that ID acts as a suppressive signal for memorization: compared to low-ID sequences, high-ID sequences are less likely to be memorized, particularly in overparameterized models and under sparse exposure. These findings highlight the interaction between scale, exposure, and complexity in shaping memorization.
title Memorization in Language Models through the Lens of Intrinsic Dimension
topic Computation and Language
url https://arxiv.org/abs/2506.09591