Does Data Contamination Detection Work (Well) for LLMs? A Survey and Evaluation on Detection Assumptions
Fuente:
arXiv
Guardado en:
| Autores principales: | Fu, Yujuan, Uzuner, Ozlem, Yetisgen, Meliha, Xia, Fei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Systematic Analysis of Large Language Models with RAG-enabled Dynamic Prompting for Medical Error Detection and Correction
por: Ahmed, Farzad, et al.
Publicado: (2025)
por: Ahmed, Farzad, et al.
Publicado: (2025)
BioMistral-NLU: Towards More Generalizable Medical Language Understanding through Instruction Tuning
por: Fu, Yujuan Velvin, et al.
Publicado: (2024)
por: Fu, Yujuan Velvin, et al.
Publicado: (2024)
A Scoping Review of Natural Language Processing in Addressing Medically Inaccurate Information: Errors, Misinformation, and Hallucination
por: Sun, Zhaoyi, et al.
Publicado: (2025)
por: Sun, Zhaoyi, et al.
Publicado: (2025)
CACER: Clinical Concept Annotations for Cancer Events and Relations
por: Fu, Yujuan, et al.
Publicado: (2024)
por: Fu, Yujuan, et al.
Publicado: (2024)
Adapting Biomedical Abstracts into Plain language using Large Language Models
por: Gangavarapu, Haritha, et al.
Publicado: (2025)
por: Gangavarapu, Haritha, et al.
Publicado: (2025)
Identifying Imaging Follow-Up in Radiology Reports: A Comparative Analysis of Traditional ML and LLM Approaches
por: Park, Namu, et al.
Publicado: (2025)
por: Park, Namu, et al.
Publicado: (2025)
MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes
por: Abacha, Asma Ben, et al.
Publicado: (2024)
por: Abacha, Asma Ben, et al.
Publicado: (2024)
Spurious Correlations and Beyond: Understanding and Mitigating Shortcut Learning in SDOH Extraction with Large Language Models
por: Sakib, Fardin Ahsan, et al.
Publicado: (2025)
por: Sakib, Fardin Ahsan, et al.
Publicado: (2025)
Extracting Social Determinants of Health from Pediatric Patient Notes Using Large Language Models: Novel Corpus and Methods
por: Fu, Yujuan, et al.
Publicado: (2024)
por: Fu, Yujuan, et al.
Publicado: (2024)
Automated Identification of Incidentalomas Requiring Follow-Up: A Multi-Anatomy Evaluation of LLM-Based and Supervised Approaches
por: Park, Namu, et al.
Publicado: (2025)
por: Park, Namu, et al.
Publicado: (2025)
A Novel Corpus of Annotated Medical Imaging Reports and Information Extraction Results Using BERT-based Language Models
por: Park, Namu, et al.
Publicado: (2024)
por: Park, Namu, et al.
Publicado: (2024)
RadTimeline: Timeline Summarization for Longitudinal Radiological Lung Findings
por: Zhou, Sitong, et al.
Publicado: (2026)
por: Zhou, Sitong, et al.
Publicado: (2026)
Prompting Large Language Models to Detect Dementia Family Caregivers
por: Biswas, Md Badsha, et al.
Publicado: (2025)
por: Biswas, Md Badsha, et al.
Publicado: (2025)
DermaVQA-DAS: Dermatology Assessment Schema (DAS) & Datasets for Closed-Ended Question Answering & Segmentation in Patient-Generated Dermatology Images
por: Yim, Wen-wai, et al.
Publicado: (2025)
por: Yim, Wen-wai, et al.
Publicado: (2025)
Contradiction to Consensus: Dual Perspective, Multi Source Retrieval Based Claim Verification with Source Level Disagreement using LLM
por: Biswas, Md Badsha, et al.
Publicado: (2026)
por: Biswas, Md Badsha, et al.
Publicado: (2026)
Large Language Model-Based Agents for Automated Research Reproducibility: An Exploratory Study in Alzheimer's Disease
por: Dobbins, Nic, et al.
Publicado: (2025)
por: Dobbins, Nic, et al.
Publicado: (2025)
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
por: Yim, Wen-wai, et al.
Publicado: (2025)
por: Yim, Wen-wai, et al.
Publicado: (2025)
DCR: Quantifying Data Contamination in LLMs Evaluation
por: Xu, Cheng, et al.
Publicado: (2025)
por: Xu, Cheng, et al.
Publicado: (2025)
Unveiling the Spectrum of Data Contamination in Language Models: A Survey from Detection to Remediation
por: Deng, Chunyuan, et al.
Publicado: (2024)
por: Deng, Chunyuan, et al.
Publicado: (2024)
UW-BioNLP at ChemoTimelines 2025: Thinking, Fine-Tuning, and Dictionary-Enhanced LLM Systems for Chemotherapy Timeline Extraction
por: Zhang, Tianmai M., et al.
Publicado: (2025)
por: Zhang, Tianmai M., et al.
Publicado: (2025)
Detecting Data Contamination in LLMs via In-Context Learning
por: Zawalski, Michał, et al.
Publicado: (2025)
por: Zawalski, Michał, et al.
Publicado: (2025)
Retrieval-Augmented Large Language Models for Schema-Constrained Clinical Information Extraction
por: Karim, A H M Rezaul, et al.
Publicado: (2026)
por: Karim, A H M Rezaul, et al.
Publicado: (2026)
Assessing Large Language Models for Structured Medical Order Extraction
por: Karim, A H M Rezaul, et al.
Publicado: (2025)
por: Karim, A H M Rezaul, et al.
Publicado: (2025)
Quantifying Data Contamination in Psychometric Evaluations of LLMs
por: Han, Jongwook, et al.
Publicado: (2025)
por: Han, Jongwook, et al.
Publicado: (2025)
Multimodal Retrieval-Augmented Generation with Large Language Models for Medical VQA
por: Karim, A H M Rezaul, et al.
Publicado: (2025)
por: Karim, A H M Rezaul, et al.
Publicado: (2025)
A Comprehensive Survey of Contamination Detection Methods in Large Language Models
por: Ravaut, Mathieu, et al.
Publicado: (2024)
por: Ravaut, Mathieu, et al.
Publicado: (2024)
CAP: Data Contamination Detection via Consistency Amplification
por: Zhao, Yi, et al.
Publicado: (2024)
por: Zhao, Yi, et al.
Publicado: (2024)
MasonTigers at SemEval-2024 Task 8: Performance Analysis of Transformer-based Models on Machine-Generated Text Detection
por: Puspo, Sadiya Sayara Chowdhury, et al.
Publicado: (2024)
por: Puspo, Sadiya Sayara Chowdhury, et al.
Publicado: (2024)
Large Language Model Capabilities in Perioperative Risk Prediction and Prognostication
por: Chung, Philip, et al.
Publicado: (2024)
por: Chung, Philip, et al.
Publicado: (2024)
A Survey on Data Contamination for Large Language Models
por: Cheng, Yuxing, et al.
Publicado: (2025)
por: Cheng, Yuxing, et al.
Publicado: (2025)
Data Contamination Quiz: A Tool to Detect and Estimate Contamination in Large Language Models
por: Golchin, Shahriar, et al.
Publicado: (2023)
por: Golchin, Shahriar, et al.
Publicado: (2023)
Benchmark Data Contamination of Large Language Models: A Survey
por: Xu, Cheng, et al.
Publicado: (2024)
por: Xu, Cheng, et al.
Publicado: (2024)
Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs
por: Balloccu, Simone, et al.
Publicado: (2024)
por: Balloccu, Simone, et al.
Publicado: (2024)
Claim Check-Worthiness Detection: How Well do LLMs Grasp Annotation Guidelines?
por: Majer, Laura, et al.
Publicado: (2024)
por: Majer, Laura, et al.
Publicado: (2024)
On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey
por: Long, Lin, et al.
Publicado: (2024)
por: Long, Lin, et al.
Publicado: (2024)
Evaluating LLMs at Detecting Errors in LLM Responses
por: Kamoi, Ryo, et al.
Publicado: (2024)
por: Kamoi, Ryo, et al.
Publicado: (2024)
Hallucination is Inevitable for LLMs with the Open World Assumption
por: Xu, Bowen
Publicado: (2025)
por: Xu, Bowen
Publicado: (2025)
Evading Data Contamination Detection for Language Models is (too) Easy
por: Dekoninck, Jasper, et al.
Publicado: (2024)
por: Dekoninck, Jasper, et al.
Publicado: (2024)
TRUCE: Private Benchmarking to Prevent Contamination and Improve Comparative Evaluation of LLMs
por: Rajore, Tanmay, et al.
Publicado: (2024)
por: Rajore, Tanmay, et al.
Publicado: (2024)
ConStat: Performance-Based Contamination Detection in Large Language Models
por: Dekoninck, Jasper, et al.
Publicado: (2024)
por: Dekoninck, Jasper, et al.
Publicado: (2024)
Ejemplares similares
-
A Systematic Analysis of Large Language Models with RAG-enabled Dynamic Prompting for Medical Error Detection and Correction
por: Ahmed, Farzad, et al.
Publicado: (2025) -
BioMistral-NLU: Towards More Generalizable Medical Language Understanding through Instruction Tuning
por: Fu, Yujuan Velvin, et al.
Publicado: (2024) -
A Scoping Review of Natural Language Processing in Addressing Medically Inaccurate Information: Errors, Misinformation, and Hallucination
por: Sun, Zhaoyi, et al.
Publicado: (2025) -
CACER: Clinical Concept Annotations for Cancer Events and Relations
por: Fu, Yujuan, et al.
Publicado: (2024) -
Adapting Biomedical Abstracts into Plain language using Large Language Models
por: Gangavarapu, Haritha, et al.
Publicado: (2025)