How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Choi, Hyeong Kyu, Khanov, Maxim, Wei, Hongxin, Li, Yixuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ARGS: Alignment as Reward-Guided Search
von: Khanov, Maxim, et al.
Veröffentlicht: (2024)
von: Khanov, Maxim, et al.
Veröffentlicht: (2024)
PICLe: Eliciting Diverse Behaviors from Large Language Models with Persona In-Context Learning
von: Choi, Hyeong Kyu, et al.
Veröffentlicht: (2024)
von: Choi, Hyeong Kyu, et al.
Veröffentlicht: (2024)
Benchmarking Benchmark Leakage in Large Language Models
von: Xu, Ruijie, et al.
Veröffentlicht: (2024)
von: Xu, Ruijie, et al.
Veröffentlicht: (2024)
Safety-Aware Fine-Tuning of Large Language Models
von: Choi, Hyeong Kyu, et al.
Veröffentlicht: (2024)
von: Choi, Hyeong Kyu, et al.
Veröffentlicht: (2024)
Preference Leakage: A Contamination Problem in LLM-as-a-judge
von: Li, Dawei, et al.
Veröffentlicht: (2025)
von: Li, Dawei, et al.
Veröffentlicht: (2025)
Detecting Distillation Data from Reasoning Models
von: Zhang, Hengxiang, et al.
Veröffentlicht: (2025)
von: Zhang, Hengxiang, et al.
Veröffentlicht: (2025)
Fine-tuning can Help Detect Pretraining Data from Large Language Models
von: Zhang, Hengxiang, et al.
Veröffentlicht: (2024)
von: Zhang, Hengxiang, et al.
Veröffentlicht: (2024)
Data Contamination Quiz: A Tool to Detect and Estimate Contamination in Large Language Models
von: Golchin, Shahriar, et al.
Veröffentlicht: (2023)
von: Golchin, Shahriar, et al.
Veröffentlicht: (2023)
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
von: Sclar, Melanie, et al.
Veröffentlicht: (2023)
von: Sclar, Melanie, et al.
Veröffentlicht: (2023)
Enhancing Chemical Reaction and Retrosynthesis Prediction with Large Language Model and Dual-task Learning
von: Lin, Xuan, et al.
Veröffentlicht: (2025)
von: Lin, Xuan, et al.
Veröffentlicht: (2025)
Recent Advances in Large Langauge Model Benchmarks against Data Contamination: From Static to Dynamic Evaluation
von: Chen, Simin, et al.
Veröffentlicht: (2025)
von: Chen, Simin, et al.
Veröffentlicht: (2025)
Detecting Data Contamination from Reinforcement Learning Post-training for Large Language Models
von: Tao, Yongding, et al.
Veröffentlicht: (2025)
von: Tao, Yongding, et al.
Veröffentlicht: (2025)
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
von: Zhao, Qihao, et al.
Veröffentlicht: (2024)
von: Zhao, Qihao, et al.
Veröffentlicht: (2024)
MemGuard: Preventing Memory Contamination in Long-Term Memory-Augmented Large Language Models
von: Ha, Hyeonjeong, et al.
Veröffentlicht: (2026)
von: Ha, Hyeonjeong, et al.
Veröffentlicht: (2026)
COPAL: Continual Pruning in Large Language Generative Models
von: Malla, Srikanth, et al.
Veröffentlicht: (2024)
von: Malla, Srikanth, et al.
Veröffentlicht: (2024)
Investigating Data Contamination for Pre-training Language Models
von: Jiang, Minhao, et al.
Veröffentlicht: (2024)
von: Jiang, Minhao, et al.
Veröffentlicht: (2024)
Should You Use Your Large Language Model to Explore or Exploit?
von: Harris, Keegan, et al.
Veröffentlicht: (2025)
von: Harris, Keegan, et al.
Veröffentlicht: (2025)
Adaptive Conformal Prediction for Improving Factuality of Generations by Large Language Models
von: Rubashevskii, Aleksandr, et al.
Veröffentlicht: (2026)
von: Rubashevskii, Aleksandr, et al.
Veröffentlicht: (2026)
FastKernels: Benchmarking GPU Kernel Generation in Production
von: Oliaro, Gabriele, et al.
Veröffentlicht: (2026)
von: Oliaro, Gabriele, et al.
Veröffentlicht: (2026)
BeyondBench: Contamination-Resistant Evaluation of Reasoning in Language Models
von: Srivastava, Gaurav, et al.
Veröffentlicht: (2025)
von: Srivastava, Gaurav, et al.
Veröffentlicht: (2025)
Time Travel in LLMs: Tracing Data Contamination in Large Language Models
von: Golchin, Shahriar, et al.
Veröffentlicht: (2023)
von: Golchin, Shahriar, et al.
Veröffentlicht: (2023)
Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications
von: Tong, Ziyi, et al.
Veröffentlicht: (2026)
von: Tong, Ziyi, et al.
Veröffentlicht: (2026)
LiveBench: A Challenging, Contamination-Limited LLM Benchmark
von: White, Colin, et al.
Veröffentlicht: (2024)
von: White, Colin, et al.
Veröffentlicht: (2024)
A Fine-tuning Dataset and Benchmark for Large Language Models for Protein Understanding
von: Shen, Yiqing, et al.
Veröffentlicht: (2024)
von: Shen, Yiqing, et al.
Veröffentlicht: (2024)
Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation
von: Baumann, Joachim, et al.
Veröffentlicht: (2025)
von: Baumann, Joachim, et al.
Veröffentlicht: (2025)
Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
von: Lan, Michael, et al.
Veröffentlicht: (2024)
von: Lan, Michael, et al.
Veröffentlicht: (2024)
Your Finetuned Large Language Model is Already a Powerful Out-of-distribution Detector
von: Zhang, Andi, et al.
Veröffentlicht: (2024)
von: Zhang, Andi, et al.
Veröffentlicht: (2024)
How Much Can We Forget about Data Contamination?
von: Bordt, Sebastian, et al.
Veröffentlicht: (2024)
von: Bordt, Sebastian, et al.
Veröffentlicht: (2024)
A Women's Health Benchmark for Large Language Models
von: Gruber, Victoria-Elisabeth, et al.
Veröffentlicht: (2025)
von: Gruber, Victoria-Elisabeth, et al.
Veröffentlicht: (2025)
Uncovering Competency Gaps in Large Language Models and Their Benchmarks
von: Bohacek, Maty, et al.
Veröffentlicht: (2025)
von: Bohacek, Maty, et al.
Veröffentlicht: (2025)
Exploring and Benchmarking the Planning Capabilities of Large Language Models
von: Bohnet, Bernd, et al.
Veröffentlicht: (2024)
von: Bohnet, Bernd, et al.
Veröffentlicht: (2024)
Reconstruct Your Previous Conversations! Comprehensively Investigating Privacy Leakage Risks in Conversations with GPT Models
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
AgentKernelArena: Generalization-Aware Benchmarking of GPU Kernel Optimization Agents
von: Younesian, Sharareh, et al.
Veröffentlicht: (2026)
von: Younesian, Sharareh, et al.
Veröffentlicht: (2026)
AfroBench: How Good are Large Language Models on African Languages?
von: Ojo, Jessica, et al.
Veröffentlicht: (2023)
von: Ojo, Jessica, et al.
Veröffentlicht: (2023)
Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States
von: Choi, Yunho, et al.
Veröffentlicht: (2026)
von: Choi, Yunho, et al.
Veröffentlicht: (2026)
The Emperor's New Clothes in Benchmarking? A Rigorous Examination of Mitigation Strategies for LLM Benchmark Data Contamination
von: Sun, Yifan, et al.
Veröffentlicht: (2025)
von: Sun, Yifan, et al.
Veröffentlicht: (2025)
NanoKnow: How to Know What Your Language Model Knows
von: Gu, Lingwei, et al.
Veröffentlicht: (2026)
von: Gu, Lingwei, et al.
Veröffentlicht: (2026)
Better Estimation of the Kullback--Leibler Divergence Between Language Models
von: Amini, Afra, et al.
Veröffentlicht: (2025)
von: Amini, Afra, et al.
Veröffentlicht: (2025)
GECOBench: A Gender-Controlled Text Dataset and Benchmark for Quantifying Biases in Explanations
von: Wilming, Rick, et al.
Veröffentlicht: (2024)
von: Wilming, Rick, et al.
Veröffentlicht: (2024)
Episodic Memories Generation and Evaluation Benchmark for Large Language Models
von: Huet, Alexis, et al.
Veröffentlicht: (2025)
von: Huet, Alexis, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ARGS: Alignment as Reward-Guided Search
von: Khanov, Maxim, et al.
Veröffentlicht: (2024) -
PICLe: Eliciting Diverse Behaviors from Large Language Models with Persona In-Context Learning
von: Choi, Hyeong Kyu, et al.
Veröffentlicht: (2024) -
Benchmarking Benchmark Leakage in Large Language Models
von: Xu, Ruijie, et al.
Veröffentlicht: (2024) -
Safety-Aware Fine-Tuning of Large Language Models
von: Choi, Hyeong Kyu, et al.
Veröffentlicht: (2024) -
Preference Leakage: A Contamination Problem in LLM-as-a-judge
von: Li, Dawei, et al.
Veröffentlicht: (2025)