The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xiong, Alexander, Zhao, Xuandong, Pappu, Aneesh, Song, Dawn
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917140268843008
author Xiong, Alexander
Zhao, Xuandong
Pappu, Aneesh
Song, Dawn
author_facet Xiong, Alexander
Zhao, Xuandong
Pappu, Aneesh
Song, Dawn
contents Large Language Models (LLMs) have demonstrated remarkable capabilities across a wide range of tasks, yet they also exhibit memorization of their training data. This phenomenon raises critical questions about model behavior, privacy risks, and the boundary between learning and memorization. Addressing these concerns, this paper synthesizes recent studies and investigates the landscape of memorization, the factors influencing it, and methods for its detection and mitigation. We explore key drivers, including training data duplication, training dynamics, and fine-tuning procedures that influence data memorization. In addition, we examine methodologies such as prefix-based extraction, membership inference, and adversarial prompting, assessing their effectiveness in detecting and measuring memorized content. Beyond technical analysis, we also explore the broader implications of memorization, including the legal and ethical implications. Finally, we discuss mitigation strategies, including data cleaning, differential privacy, and post-training unlearning, while highlighting open challenges in balancing the need to minimize harmful memorization with model utility. This paper provides a comprehensive overview of the current state of research on LLM memorization across technical, privacy, and performance dimensions, identifying critical directions for future work.
format Preprint
id arxiv_https___arxiv_org_abs_2507_05578
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation
Xiong, Alexander
Zhao, Xuandong
Pappu, Aneesh
Song, Dawn
Machine Learning
Computation and Language
Cryptography and Security
Large Language Models (LLMs) have demonstrated remarkable capabilities across a wide range of tasks, yet they also exhibit memorization of their training data. This phenomenon raises critical questions about model behavior, privacy risks, and the boundary between learning and memorization. Addressing these concerns, this paper synthesizes recent studies and investigates the landscape of memorization, the factors influencing it, and methods for its detection and mitigation. We explore key drivers, including training data duplication, training dynamics, and fine-tuning procedures that influence data memorization. In addition, we examine methodologies such as prefix-based extraction, membership inference, and adversarial prompting, assessing their effectiveness in detecting and measuring memorized content. Beyond technical analysis, we also explore the broader implications of memorization, including the legal and ethical implications. Finally, we discuss mitigation strategies, including data cleaning, differential privacy, and post-training unlearning, while highlighting open challenges in balancing the need to minimize harmful memorization with model utility. This paper provides a comprehensive overview of the current state of research on LLM memorization across technical, privacy, and performance dimensions, identifying critical directions for future work.
title The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation
topic Machine Learning
Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2507.05578