Detecting Non-Membership in LLM Training Data via Rank Correlations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shetty, Pranav, Haque, Mirazul, Ma, Zhiqiang, Liu, Xiaomo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908916024082432
author Shetty, Pranav
Haque, Mirazul
Ma, Zhiqiang
Liu, Xiaomo
author_facet Shetty, Pranav
Haque, Mirazul
Ma, Zhiqiang
Liu, Xiaomo
contents As large language models (LLMs) are trained on increasingly vast and opaque text corpora, determining which data contributed to training has become essential for copyright enforcement, compliance auditing, and user trust. While prior work focuses on detecting whether a dataset was used in training (membership inference), the complementary problem -- verifying that a dataset was not used -- has received little attention. We address this gap by introducing PRISM, a test that detects dataset-level non-membership using only grey-box access to model logits. Our key insight is that two models that have not seen a dataset exhibit higher rank correlation in their normalized token log probabilities than when one model has been trained on that data. Using this observation, we construct a correlation-based test that detects non-membership. Empirically, PRISM reliably rules out membership in training data across all datasets tested while avoiding false positives, thus offering a framework for verifying that specific datasets were excluded from LLM training.
format Preprint
id arxiv_https___arxiv_org_abs_2603_22707
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Detecting Non-Membership in LLM Training Data via Rank Correlations
Shetty, Pranav
Haque, Mirazul
Ma, Zhiqiang
Liu, Xiaomo
Computation and Language
As large language models (LLMs) are trained on increasingly vast and opaque text corpora, determining which data contributed to training has become essential for copyright enforcement, compliance auditing, and user trust. While prior work focuses on detecting whether a dataset was used in training (membership inference), the complementary problem -- verifying that a dataset was not used -- has received little attention. We address this gap by introducing PRISM, a test that detects dataset-level non-membership using only grey-box access to model logits. Our key insight is that two models that have not seen a dataset exhibit higher rank correlation in their normalized token log probabilities than when one model has been trained on that data. Using this observation, we construct a correlation-based test that detects non-membership. Empirically, PRISM reliably rules out membership in training data across all datasets tested while avoiding false positives, thus offering a framework for verifying that specific datasets were excluded from LLM training.
title Detecting Non-Membership in LLM Training Data via Rank Correlations
topic Computation and Language
url https://arxiv.org/abs/2603.22707