Saved in:
Bibliographic Details
Main Authors: Chen, Simin, Chen, Yiming, Li, Zexin, Jiang, Yifan, Wan, Zhongwei, He, Yixin, Ran, Dezhi, Gu, Tianle, Li, Haizhou, Xie, Tao, Ray, Baishakhi
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2502.17521
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909815748427776
author Chen, Simin
Chen, Yiming
Li, Zexin
Jiang, Yifan
Wan, Zhongwei
He, Yixin
Ran, Dezhi
Gu, Tianle
Li, Haizhou
Xie, Tao
Ray, Baishakhi
author_facet Chen, Simin
Chen, Yiming
Li, Zexin
Jiang, Yifan
Wan, Zhongwei
He, Yixin
Ran, Dezhi
Gu, Tianle
Li, Haizhou
Xie, Tao
Ray, Baishakhi
contents Data contamination has received increasing attention in the era of large language models (LLMs) due to their reliance on vast Internet-derived training corpora. To mitigate the risk of potential data contamination, LLM benchmarking has undergone a transformation from static to dynamic benchmarking. In this work, we conduct an in-depth analysis of existing static to dynamic benchmarking methods aimed at reducing data contamination risks. We first examine methods that enhance static benchmarks and identify their inherent limitations. We then highlight a critical gap-the lack of standardized criteria for evaluating dynamic benchmarks. Based on this observation, we propose a series of optimal design principles for dynamic benchmarking and analyze the limitations of existing dynamic benchmarks. This survey provides a concise yet comprehensive overview of recent advancements in data contamination research, offering valuable insights and a clear guide for future research efforts. We maintain a GitHub repository to continuously collect both static and dynamic benchmarking methods for LLMs. The repository can be found at this link.
format Preprint
id arxiv_https___arxiv_org_abs_2502_17521
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Recent Advances in Large Langauge Model Benchmarks against Data Contamination: From Static to Dynamic Evaluation
Chen, Simin
Chen, Yiming
Li, Zexin
Jiang, Yifan
Wan, Zhongwei
He, Yixin
Ran, Dezhi
Gu, Tianle
Li, Haizhou
Xie, Tao
Ray, Baishakhi
Machine Learning
Artificial Intelligence
Computation and Language
Data contamination has received increasing attention in the era of large language models (LLMs) due to their reliance on vast Internet-derived training corpora. To mitigate the risk of potential data contamination, LLM benchmarking has undergone a transformation from static to dynamic benchmarking. In this work, we conduct an in-depth analysis of existing static to dynamic benchmarking methods aimed at reducing data contamination risks. We first examine methods that enhance static benchmarks and identify their inherent limitations. We then highlight a critical gap-the lack of standardized criteria for evaluating dynamic benchmarks. Based on this observation, we propose a series of optimal design principles for dynamic benchmarking and analyze the limitations of existing dynamic benchmarks. This survey provides a concise yet comprehensive overview of recent advancements in data contamination research, offering valuable insights and a clear guide for future research efforts. We maintain a GitHub repository to continuously collect both static and dynamic benchmarking methods for LLMs. The repository can be found at this link.
title Recent Advances in Large Langauge Model Benchmarks against Data Contamination: From Static to Dynamic Evaluation
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2502.17521