Detecting Data Contamination in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Janicki, Juliusz, Chamezopoulos, Savvas, Kanoulas, Evangelos, Tsatsaronis, Georgios |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimizing Numerical Estimation and Operational Efficiency in the Legal Domain through Large Language Models
by: Huang, Jia-Hong, et al.
Published: (2024)
by: Huang, Jia-Hong, et al.
Published: (2024)
Data Contamination Quiz: A Tool to Detect and Estimate Contamination in Large Language Models
by: Golchin, Shahriar, et al.
Published: (2023)
by: Golchin, Shahriar, et al.
Published: (2023)
A Survey on Recent Advances in Conversational Data Generation
by: Soudani, Heydar, et al.
Published: (2024)
by: Soudani, Heydar, et al.
Published: (2024)
Knowledge-Enhanced Conversational Recommendation via Transformer-based Sequential Modelling
by: Zou, Jie, et al.
Published: (2024)
by: Zou, Jie, et al.
Published: (2024)
Spectral Tempering for Embedding Compression in Dense Passage Retrieval
by: Li, Yongkang, et al.
Published: (2026)
by: Li, Yongkang, et al.
Published: (2026)
Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation
by: Liu, Ming, et al.
Published: (2025)
by: Liu, Ming, et al.
Published: (2025)
Trustworthy AI: Securing Sensitive Data in Large Language Models
by: Feretzakis, Georgios, et al.
Published: (2024)
by: Feretzakis, Georgios, et al.
Published: (2024)
Towards Data Contamination Detection for Modern Large Language Models: Limitations, Inconsistencies, and Oracle Challenges
by: Samuel, Vinay, et al.
Published: (2024)
by: Samuel, Vinay, et al.
Published: (2024)
An Open Source Data Contamination Report for Large Language Models
by: Li, Yucheng, et al.
Published: (2023)
by: Li, Yucheng, et al.
Published: (2023)
Investigating Data Contamination in Modern Benchmarks for Large Language Models
by: Deng, Chunyuan, et al.
Published: (2023)
by: Deng, Chunyuan, et al.
Published: (2023)
SubSearch: Intermediate Rewards for Unsupervised Guided Reasoning in Complex Retrieval
by: Petcu, Roxana, et al.
Published: (2026)
by: Petcu, Roxana, et al.
Published: (2026)
Evaluation of Large Language Models for Anomaly Detection in Autonomous Vehicles
by: Loukas, Petros, et al.
Published: (2025)
by: Loukas, Petros, et al.
Published: (2025)
Detecting Data Contamination from Reinforcement Learning Post-training for Large Language Models
by: Tao, Yongding, et al.
Published: (2025)
by: Tao, Yongding, et al.
Published: (2025)
DVD: A Robust Method for Detecting Variant Contamination in Large Language Model Evaluation
by: Liang, Renzhao, et al.
Published: (2026)
by: Liang, Renzhao, et al.
Published: (2026)
A Proof-of-Concept for Explainable Disease Diagnosis Using Large Language Models and Answer Set Programming
by: Gemou, Ioanna, et al.
Published: (2025)
by: Gemou, Ioanna, et al.
Published: (2025)
Evading Data Contamination Detection for Language Models is (too) Easy
by: Dekoninck, Jasper, et al.
Published: (2024)
by: Dekoninck, Jasper, et al.
Published: (2024)
Learning to Ask: Conversational Product Search via Representation Learning
by: Zou, Jie, et al.
Published: (2024)
by: Zou, Jie, et al.
Published: (2024)
AdEval: Alignment-based Dynamic Evaluation to Mitigate Data Contamination in Large Language Models
by: Fan, Yang
Published: (2025)
by: Fan, Yang
Published: (2025)
Gradient Weight-normalized Low-rank Projection for Efficient LLM Training
by: Huang, Jia-Hong, et al.
Published: (2024)
by: Huang, Jia-Hong, et al.
Published: (2024)
Time Travel in LLMs: Tracing Data Contamination in Large Language Models
by: Golchin, Shahriar, et al.
Published: (2023)
by: Golchin, Shahriar, et al.
Published: (2023)
PaCoST: Paired Confidence Significance Testing for Benchmark Contamination Detection in Large Language Models
by: Zhang, Huixuan, et al.
Published: (2024)
by: Zhang, Huixuan, et al.
Published: (2024)
Sensitivity of Small Language Models to Fine-tuning Data Contamination
by: Scaria, Nicy, et al.
Published: (2025)
by: Scaria, Nicy, et al.
Published: (2025)
A Novel Automatic Framework for Speaker Drift Detection in Synthesized Speech
by: Huang, Jia-Hong, et al.
Published: (2026)
by: Huang, Jia-Hong, et al.
Published: (2026)
Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination
by: Chen, Simin, et al.
Published: (2025)
by: Chen, Simin, et al.
Published: (2025)
On the Complexity of Winner Determination and Strategic Control in Conditional Approval Voting
by: Markakis, Evangelos, et al.
Published: (2022)
by: Markakis, Evangelos, et al.
Published: (2022)
QFMTS: Generating Query-Focused Summaries over Multi-Table Inputs
by: Zhang, Weijia, et al.
Published: (2024)
by: Zhang, Weijia, et al.
Published: (2024)
Investigating Data Contamination for Pre-training Language Models
by: Jiang, Minhao, et al.
Published: (2024)
by: Jiang, Minhao, et al.
Published: (2024)
VLM-RRT: Vision Language Model Guided RRT Search for Autonomous UAV Navigation
by: Ye, Jianlin, et al.
Published: (2025)
by: Ye, Jianlin, et al.
Published: (2025)
Data Contamination Can Cross Language Barriers
by: Yao, Feng, et al.
Published: (2024)
by: Yao, Feng, et al.
Published: (2024)
Generalization or Memorization: Data Contamination and Trustworthy Evaluation for Large Language Models
by: Dong, Yihong, et al.
Published: (2024)
by: Dong, Yihong, et al.
Published: (2024)
Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications
by: Tong, Ziyi, et al.
Published: (2026)
by: Tong, Ziyi, et al.
Published: (2026)
Query Decomposition for RAG: Balancing Exploration-Exploitation
by: Petcu, Roxana, et al.
Published: (2025)
by: Petcu, Roxana, et al.
Published: (2025)
Improving the Robustness of Dense Retrievers Against Typos via Multi-Positive Contrastive Learning
by: Sidiropoulos, Georgios, et al.
Published: (2024)
by: Sidiropoulos, Georgios, et al.
Published: (2024)
A Multimodal Dense Retrieval Approach for Speech-Based Open-Domain Question Answering
by: Sidiropoulos, Georgios, et al.
Published: (2024)
by: Sidiropoulos, Georgios, et al.
Published: (2024)
Categorical semantics of compositional reinforcement learning
by: Bakirtzis, Georgios, et al.
Published: (2022)
by: Bakirtzis, Georgios, et al.
Published: (2022)
Deep Positive-Unlabeled Anomaly Detection for Contaminated Unlabeled Data
by: Takahashi, Hiroshi, et al.
Published: (2024)
by: Takahashi, Hiroshi, et al.
Published: (2024)
RADAR: Mechanistic Pathways for Detecting Data Contamination in LLM Evaluation
by: Kattamuri, Ashish, et al.
Published: (2025)
by: Kattamuri, Ashish, et al.
Published: (2025)
Anomaly Detection with Adaptive and Aggressive Rejection for Contaminated Training Data
by: Lee, Jungi, et al.
Published: (2025)
by: Lee, Jungi, et al.
Published: (2025)
The Heap: A Contamination-Free Multilingual Code Dataset for Evaluating Large Language Models
by: Katzy, Jonathan, et al.
Published: (2025)
by: Katzy, Jonathan, et al.
Published: (2025)
Beyond Surface-Level Similarity: Hierarchical Contamination Detection for Synthetic Training Data in Foundation Models
by: Mehta, Sushant
Published: (2025)
by: Mehta, Sushant
Published: (2025)
Similar Items
-
Optimizing Numerical Estimation and Operational Efficiency in the Legal Domain through Large Language Models
by: Huang, Jia-Hong, et al.
Published: (2024) -
Data Contamination Quiz: A Tool to Detect and Estimate Contamination in Large Language Models
by: Golchin, Shahriar, et al.
Published: (2023) -
A Survey on Recent Advances in Conversational Data Generation
by: Soudani, Heydar, et al.
Published: (2024) -
Knowledge-Enhanced Conversational Recommendation via Transformer-based Sequential Modelling
by: Zou, Jie, et al.
Published: (2024) -
Spectral Tempering for Embedding Compression in Dense Passage Retrieval
by: Li, Yongkang, et al.
Published: (2026)