Language Models May Verbatim Complete Text They Were Not Explicitly Trained On
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910891894636544 |
|---|---|
| author | Liu, Ken Ziyu Choquette-Choo, Christopher A. Jagielski, Matthew Kairouz, Peter Koyejo, Sanmi Liang, Percy Papernot, Nicolas |
| author_facet | Liu, Ken Ziyu Choquette-Choo, Christopher A. Jagielski, Matthew Kairouz, Peter Koyejo, Sanmi Liang, Percy Papernot, Nicolas |
| contents | An important question today is whether a given text was used to train a large language model (LLM). A \emph{completion} test is often employed: check if the LLM completes a sufficiently complex text. This, however, requires a ground-truth definition of membership; most commonly, it is defined as a member based on the $n$-gram overlap between the target text and any text in the dataset. In this work, we demonstrate that this $n$-gram based membership definition can be effectively gamed. We study scenarios where sequences are \emph{non-members} for a given $n$ and we find that completion tests still succeed. We find many natural cases of this phenomenon by retraining LLMs from scratch after removing all training samples that were completed; these cases include exact duplicates, near-duplicates, and even short overlaps. They showcase that it is difficult to find a single viable choice of $n$ for membership definitions. Using these insights, we design adversarial datasets that can cause a given target sequence to be completed without containing it, for any reasonable choice of $n$. Our findings highlight the inadequacy of $n$-gram membership, suggesting membership definitions fail to account for auxiliary information available to the training algorithm. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_17514 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Language Models May Verbatim Complete Text They Were Not Explicitly Trained On Liu, Ken Ziyu Choquette-Choo, Christopher A. Jagielski, Matthew Kairouz, Peter Koyejo, Sanmi Liang, Percy Papernot, Nicolas Computation and Language Artificial Intelligence Cryptography and Security Machine Learning An important question today is whether a given text was used to train a large language model (LLM). A \emph{completion} test is often employed: check if the LLM completes a sufficiently complex text. This, however, requires a ground-truth definition of membership; most commonly, it is defined as a member based on the $n$-gram overlap between the target text and any text in the dataset. In this work, we demonstrate that this $n$-gram based membership definition can be effectively gamed. We study scenarios where sequences are \emph{non-members} for a given $n$ and we find that completion tests still succeed. We find many natural cases of this phenomenon by retraining LLMs from scratch after removing all training samples that were completed; these cases include exact duplicates, near-duplicates, and even short overlaps. They showcase that it is difficult to find a single viable choice of $n$ for membership definitions. Using these insights, we design adversarial datasets that can cause a given target sequence to be completed without containing it, for any reasonable choice of $n$. Our findings highlight the inadequacy of $n$-gram membership, suggesting membership definitions fail to account for auxiliary information available to the training algorithm. |
| title | Language Models May Verbatim Complete Text They Were Not Explicitly Trained On |
| topic | Computation and Language Artificial Intelligence Cryptography and Security Machine Learning |
| url | https://arxiv.org/abs/2503.17514 |