Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Balloccu, Simone, Schmidtová, Patrícia, Lango, Mateusz, Dušek, Ondřej |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM Agents Implement an NLG System from Scratch: Building Interpretable Rule-Based RDF-to-Text Generators
by: Lango, Mateusz, et al.
Published: (2025)
by: Lango, Mateusz, et al.
Published: (2025)
Leveraging Large Language Models for Building Interpretable Rule-Based Data-to-Text Systems
by: Warczyński, Jędrzej, et al.
Published: (2025)
by: Warczyński, Jędrzej, et al.
Published: (2025)
Faithful and Plausible Natural Language Explanations for Image Classification: A Pipeline Approach
by: Wojciechowski, Adam, et al.
Published: (2024)
by: Wojciechowski, Adam, et al.
Published: (2024)
factgenie: A Framework for Span-based Evaluation of Generated Texts
by: Kasner, Zdeněk, et al.
Published: (2024)
by: Kasner, Zdeněk, et al.
Published: (2024)
OpeNLGauge: An Explainable Metric for NLG Evaluation with Open-Weights LLMs
by: Kartáč, Ivan, et al.
Published: (2025)
by: Kartáč, Ivan, et al.
Published: (2025)
Reasoning Gets Harder for LLMs Inside A Dialogue
by: Kartáč, Ivan, et al.
Published: (2026)
by: Kartáč, Ivan, et al.
Published: (2026)
When LLMs Can't Help: Real-World Evaluation of LLMs in Nutrition
by: Li, Karen Jia-Hui, et al.
Published: (2025)
by: Li, Karen Jia-Hui, et al.
Published: (2025)
Real-World Summarization: When Evaluation Reaches Its Limits
by: Schmidtová, Patrícia, et al.
Published: (2025)
by: Schmidtová, Patrícia, et al.
Published: (2025)
LLMs as Span Annotators: A Comparative Study of LLMs and Humans
by: Kasner, Zdeněk, et al.
Published: (2025)
by: Kasner, Zdeněk, et al.
Published: (2025)
SRS-Stories: Vocabulary-constrained multilingual story generation for language learning
by: Kamzela, Wiktor, et al.
Published: (2025)
by: Kamzela, Wiktor, et al.
Published: (2025)
Automatic Metrics in Natural Language Generation: A Survey of Current Evaluation Practices
by: Schmidtová, Patrícia, et al.
Published: (2024)
by: Schmidtová, Patrícia, et al.
Published: (2024)
Polish-ASTE: Aspect-Sentiment Triplet Extraction Datasets for Polish
by: Lango, Marta, et al.
Published: (2025)
by: Lango, Marta, et al.
Published: (2025)
A Survey of Text Style Transfer: Applications and Ethical Implications
by: Mukherjee, Sourabrata, et al.
Published: (2024)
by: Mukherjee, Sourabrata, et al.
Published: (2024)
ASTE Transformer Modelling Dependencies in Aspect-Sentiment Triplet Extraction
by: Naglik, Iwo, et al.
Published: (2024)
by: Naglik, Iwo, et al.
Published: (2024)
Hypothesis-Driven Feature Manifold Analysis in LLMs via Supervised Multi-Dimensional Scaling
by: Tiblias, Federico, et al.
Published: (2025)
by: Tiblias, Federico, et al.
Published: (2025)
The Problem of Coherence in Natural Language Explanations of Recommendations
by: Raczyński, Jakub, et al.
Published: (2023)
by: Raczyński, Jakub, et al.
Published: (2023)
FreshTab: Sourcing Fresh Data for Table-to-Text Generation Evaluation
by: Onderková, Kristýna, et al.
Published: (2025)
by: Onderková, Kristýna, et al.
Published: (2025)
An Open Source Data Contamination Report for Large Language Models
by: Li, Yucheng, et al.
Published: (2023)
by: Li, Yucheng, et al.
Published: (2023)
UFAL-CUNI at SemEval-2026 Task 11: An Efficient Modular Neuro-symbolic Method for Syllogistic Reasoning
by: Kartáč, Ivan, et al.
Published: (2026)
by: Kartáč, Ivan, et al.
Published: (2026)
Both Text and Images Leaked! A Systematic Analysis of Data Contamination in Multimodal LLM
by: Song, Dingjie, et al.
Published: (2024)
by: Song, Dingjie, et al.
Published: (2024)
When Benchmarks Leak: Inference-Time Decontamination for LLMs
by: Chai, Jianzhe, et al.
Published: (2026)
by: Chai, Jianzhe, et al.
Published: (2026)
Beyond Traditional Benchmarks: Analyzing Behaviors of Open LLMs on Data-to-Text Generation
by: Kasner, Zdeněk, et al.
Published: (2024)
by: Kasner, Zdeněk, et al.
Published: (2024)
Leaking LoRa: An Evaluation of Password Leaks and Knowledge Storage in Large Language Models
by: Marinelli, Ryan, et al.
Published: (2025)
by: Marinelli, Ryan, et al.
Published: (2025)
Exploring ReAct Prompting for Task-Oriented Dialogue: Insights and Shortcomings
by: Elizabeth, Michelle, et al.
Published: (2024)
by: Elizabeth, Michelle, et al.
Published: (2024)
AnimatedLLM: Explaining LLMs with Interactive Visualizations
by: Kasner, Zdeněk, et al.
Published: (2025)
by: Kasner, Zdeněk, et al.
Published: (2025)
Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
by: Papi, Sara, et al.
Published: (2025)
by: Papi, Sara, et al.
Published: (2025)
Detecting Data Contamination in LLMs via In-Context Learning
by: Zawalski, Michał, et al.
Published: (2025)
by: Zawalski, Michał, et al.
Published: (2025)
Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort
by: Wang, Xinpeng, et al.
Published: (2025)
by: Wang, Xinpeng, et al.
Published: (2025)
CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
AdEval: Alignment-based Dynamic Evaluation to Mitigate Data Contamination in Large Language Models
by: Fan, Yang
Published: (2025)
by: Fan, Yang
Published: (2025)
LessLeak-Bench: A First Investigation of Data Leakage in LLMs Across 83 Software Engineering Benchmarks
by: Zhou, Xin, et al.
Published: (2025)
by: Zhou, Xin, et al.
Published: (2025)
Data Contamination Can Cross Language Barriers
by: Yao, Feng, et al.
Published: (2024)
by: Yao, Feng, et al.
Published: (2024)
Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation
by: Mahdavi, Sadegh, et al.
Published: (2025)
by: Mahdavi, Sadegh, et al.
Published: (2025)
Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation
by: Kocyigit, Muhammed Yusuf, et al.
Published: (2025)
by: Kocyigit, Muhammed Yusuf, et al.
Published: (2025)
On the Importance and Evaluation of Narrativity in Natural Language AI Explanations
by: Cedro, Mateusz, et al.
Published: (2026)
by: Cedro, Mateusz, et al.
Published: (2026)
Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs
by: Kabir, Mohsinul, et al.
Published: (2025)
by: Kabir, Mohsinul, et al.
Published: (2025)
LatestEval: Addressing Data Contamination in Language Model Evaluation through Dynamic and Time-Sensitive Test Construction
by: Li, Yucheng, et al.
Published: (2023)
by: Li, Yucheng, et al.
Published: (2023)
LLMs in Interpreting Legal Documents
by: Corbo, Simone
Published: (2025)
by: Corbo, Simone
Published: (2025)
Sensitivity of Small Language Models to Fine-tuning Data Contamination
by: Scaria, Nicy, et al.
Published: (2025)
by: Scaria, Nicy, et al.
Published: (2025)
Investigating Data Contamination in Modern Benchmarks for Large Language Models
by: Deng, Chunyuan, et al.
Published: (2023)
by: Deng, Chunyuan, et al.
Published: (2023)
Similar Items
-
LLM Agents Implement an NLG System from Scratch: Building Interpretable Rule-Based RDF-to-Text Generators
by: Lango, Mateusz, et al.
Published: (2025) -
Leveraging Large Language Models for Building Interpretable Rule-Based Data-to-Text Systems
by: Warczyński, Jędrzej, et al.
Published: (2025) -
Faithful and Plausible Natural Language Explanations for Image Classification: A Pipeline Approach
by: Wojciechowski, Adam, et al.
Published: (2024) -
factgenie: A Framework for Span-based Evaluation of Generated Texts
by: Kasner, Zdeněk, et al.
Published: (2024) -
OpeNLGauge: An Explainable Metric for NLG Evaluation with Open-Weights LLMs
by: Kartáč, Ivan, et al.
Published: (2025)