Detecting Data Contamination in LLMs via In-Context Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zawalski, Michał, Boubdir, Meriem, Bałazy, Klaudia, Nushi, Besmira, Ribalta, Pablo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913117920821248
author Zawalski, Michał
Boubdir, Meriem
Bałazy, Klaudia
Nushi, Besmira
Ribalta, Pablo
author_facet Zawalski, Michał
Boubdir, Meriem
Bałazy, Klaudia
Nushi, Besmira
Ribalta, Pablo
contents We present Contamination Detection via Context (CoDeC), a practical and accurate method to detect and quantify training data contamination in large language models. CoDeC distinguishes between data memorized during training and data outside the training distribution by measuring how in-context learning affects model performance. We find that in-context examples typically boost confidence for unseen datasets but may reduce it when the dataset was part of training, due to disrupted memorization patterns. Experiments show that CoDeC produces interpretable contamination scores that clearly separate seen and unseen datasets, and reveals strong evidence of memorization in open-weight models with undisclosed training corpora. The method is simple, automated, and both model- and dataset-agnostic, making it easy to integrate with benchmark evaluations.
format Preprint
id arxiv_https___arxiv_org_abs_2510_27055
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Detecting Data Contamination in LLMs via In-Context Learning
Zawalski, Michał
Boubdir, Meriem
Bałazy, Klaudia
Nushi, Besmira
Ribalta, Pablo
Computation and Language
Artificial Intelligence
I.2.7
We present Contamination Detection via Context (CoDeC), a practical and accurate method to detect and quantify training data contamination in large language models. CoDeC distinguishes between data memorized during training and data outside the training distribution by measuring how in-context learning affects model performance. We find that in-context examples typically boost confidence for unseen datasets but may reduce it when the dataset was part of training, due to disrupted memorization patterns. Experiments show that CoDeC produces interpretable contamination scores that clearly separate seen and unseen datasets, and reveals strong evidence of memorization in open-weight models with undisclosed training corpora. The method is simple, automated, and both model- and dataset-agnostic, making it easy to integrate with benchmark evaluations.
title Detecting Data Contamination in LLMs via In-Context Learning
topic Computation and Language
Artificial Intelligence
I.2.7
url https://arxiv.org/abs/2510.27055