Language Models May Verbatim Complete Text They Were Not Explicitly Trained On

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Ken Ziyu, Choquette-Choo, Christopher A., Jagielski, Matthew, Kairouz, Peter, Koyejo, Sanmi, Liang, Percy, Papernot, Nicolas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910891894636544
author Liu, Ken Ziyu
Choquette-Choo, Christopher A.
Jagielski, Matthew
Kairouz, Peter
Koyejo, Sanmi
Liang, Percy
Papernot, Nicolas
author_facet Liu, Ken Ziyu
Choquette-Choo, Christopher A.
Jagielski, Matthew
Kairouz, Peter
Koyejo, Sanmi
Liang, Percy
Papernot, Nicolas
contents An important question today is whether a given text was used to train a large language model (LLM). A \emph{completion} test is often employed: check if the LLM completes a sufficiently complex text. This, however, requires a ground-truth definition of membership; most commonly, it is defined as a member based on the $n$-gram overlap between the target text and any text in the dataset. In this work, we demonstrate that this $n$-gram based membership definition can be effectively gamed. We study scenarios where sequences are \emph{non-members} for a given $n$ and we find that completion tests still succeed. We find many natural cases of this phenomenon by retraining LLMs from scratch after removing all training samples that were completed; these cases include exact duplicates, near-duplicates, and even short overlaps. They showcase that it is difficult to find a single viable choice of $n$ for membership definitions. Using these insights, we design adversarial datasets that can cause a given target sequence to be completed without containing it, for any reasonable choice of $n$. Our findings highlight the inadequacy of $n$-gram membership, suggesting membership definitions fail to account for auxiliary information available to the training algorithm.
format Preprint
id arxiv_https___arxiv_org_abs_2503_17514
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Language Models May Verbatim Complete Text They Were Not Explicitly Trained On
Liu, Ken Ziyu
Choquette-Choo, Christopher A.
Jagielski, Matthew
Kairouz, Peter
Koyejo, Sanmi
Liang, Percy
Papernot, Nicolas
Computation and Language
Artificial Intelligence
Cryptography and Security
Machine Learning
An important question today is whether a given text was used to train a large language model (LLM). A \emph{completion} test is often employed: check if the LLM completes a sufficiently complex text. This, however, requires a ground-truth definition of membership; most commonly, it is defined as a member based on the $n$-gram overlap between the target text and any text in the dataset. In this work, we demonstrate that this $n$-gram based membership definition can be effectively gamed. We study scenarios where sequences are \emph{non-members} for a given $n$ and we find that completion tests still succeed. We find many natural cases of this phenomenon by retraining LLMs from scratch after removing all training samples that were completed; these cases include exact duplicates, near-duplicates, and even short overlaps. They showcase that it is difficult to find a single viable choice of $n$ for membership definitions. Using these insights, we design adversarial datasets that can cause a given target sequence to be completed without containing it, for any reasonable choice of $n$. Our findings highlight the inadequacy of $n$-gram membership, suggesting membership definitions fail to account for auxiliary information available to the training algorithm.
title Language Models May Verbatim Complete Text They Were Not Explicitly Trained On
topic Computation and Language
Artificial Intelligence
Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2503.17514