Investigating Data Contamination for Pre-training Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Minhao, Liu, Ken Ziyu, Zhong, Ming, Schaeffer, Rylan, Ouyang, Siru, Han, Jiawei, Koyejo, Sanmi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910294710681600
author Jiang, Minhao
Liu, Ken Ziyu
Zhong, Ming
Schaeffer, Rylan
Ouyang, Siru
Han, Jiawei
Koyejo, Sanmi
author_facet Jiang, Minhao
Liu, Ken Ziyu
Zhong, Ming
Schaeffer, Rylan
Ouyang, Siru
Han, Jiawei
Koyejo, Sanmi
contents Language models pre-trained on web-scale corpora demonstrate impressive capabilities on diverse downstream tasks. However, there is increasing concern whether such capabilities might arise from evaluation datasets being included in the pre-training corpus -- a phenomenon known as \textit{data contamination} -- in a manner that artificially increases performance. There has been little understanding of how this potential contamination might influence LMs' performance on downstream tasks. In this paper, we explore the impact of data contamination at the pre-training stage by pre-training a series of GPT-2 models \textit{from scratch}. We highlight the effect of both text contamination (\textit{i.e.}\ input text of the evaluation samples) and ground-truth contamination (\textit{i.e.}\ the prompts asked on the input and the desired outputs) from evaluation data. We also investigate the effects of repeating contamination for various downstream tasks. Additionally, we examine the prevailing n-gram-based definitions of contamination within current LLM reports, pinpointing their limitations and inadequacy. Our findings offer new insights into data contamination's effects on language model capabilities and underscore the need for independent, comprehensive contamination assessments in LLM studies.
format Preprint
id arxiv_https___arxiv_org_abs_2401_06059
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Investigating Data Contamination for Pre-training Language Models
Jiang, Minhao
Liu, Ken Ziyu
Zhong, Ming
Schaeffer, Rylan
Ouyang, Siru
Han, Jiawei
Koyejo, Sanmi
Computation and Language
Artificial Intelligence
Machine Learning
Language models pre-trained on web-scale corpora demonstrate impressive capabilities on diverse downstream tasks. However, there is increasing concern whether such capabilities might arise from evaluation datasets being included in the pre-training corpus -- a phenomenon known as \textit{data contamination} -- in a manner that artificially increases performance. There has been little understanding of how this potential contamination might influence LMs' performance on downstream tasks. In this paper, we explore the impact of data contamination at the pre-training stage by pre-training a series of GPT-2 models \textit{from scratch}. We highlight the effect of both text contamination (\textit{i.e.}\ input text of the evaluation samples) and ground-truth contamination (\textit{i.e.}\ the prompts asked on the input and the desired outputs) from evaluation data. We also investigate the effects of repeating contamination for various downstream tasks. Additionally, we examine the prevailing n-gram-based definitions of contamination within current LLM reports, pinpointing their limitations and inadequacy. Our findings offer new insights into data contamination's effects on language model capabilities and underscore the need for independent, comprehensive contamination assessments in LLM studies.
title Investigating Data Contamination for Pre-training Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2401.06059