CAP: Data Contamination Detection via Consistency Amplification
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Yi, Li, Jing, Yang, Linyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability Curvature
von: Bao, Guangsheng, et al.
Veröffentlicht: (2023)
von: Bao, Guangsheng, et al.
Veröffentlicht: (2023)
LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories
von: He, Zirui, et al.
Veröffentlicht: (2025)
von: He, Zirui, et al.
Veröffentlicht: (2025)
Proposal Report for the 2nd SciCAP Competition 2024
von: Li, Pengpeng, et al.
Veröffentlicht: (2024)
von: Li, Pengpeng, et al.
Veröffentlicht: (2024)
Unveiling the Spectrum of Data Contamination in Language Models: A Survey from Detection to Remediation
von: Deng, Chunyuan, et al.
Veröffentlicht: (2024)
von: Deng, Chunyuan, et al.
Veröffentlicht: (2024)
A Survey on Data Contamination for Large Language Models
von: Cheng, Yuxing, et al.
Veröffentlicht: (2025)
von: Cheng, Yuxing, et al.
Veröffentlicht: (2025)
Confabulations from ACL Publications (CAP): A Dataset for Scientific Hallucination Detection
von: Gamba, Federica, et al.
Veröffentlicht: (2025)
von: Gamba, Federica, et al.
Veröffentlicht: (2025)
Detecting Data Contamination in LLMs via In-Context Learning
von: Zawalski, Michał, et al.
Veröffentlicht: (2025)
von: Zawalski, Michał, et al.
Veröffentlicht: (2025)
Does Data Contamination Detection Work (Well) for LLMs? A Survey and Evaluation on Detection Assumptions
von: Fu, Yujuan, et al.
Veröffentlicht: (2024)
von: Fu, Yujuan, et al.
Veröffentlicht: (2024)
MCEval: A Dynamic Framework for Fair Multilingual Cultural Evaluation of LLMs
von: Huang, Shulin, et al.
Veröffentlicht: (2025)
von: Huang, Shulin, et al.
Veröffentlicht: (2025)
Data Contamination Quiz: A Tool to Detect and Estimate Contamination in Large Language Models
von: Golchin, Shahriar, et al.
Veröffentlicht: (2023)
von: Golchin, Shahriar, et al.
Veröffentlicht: (2023)
A Rationale-centric Counterfactual Data Augmentation Method for Cross-Document Event Coreference Resolution
von: Ding, Bowen, et al.
Veröffentlicht: (2024)
von: Ding, Bowen, et al.
Veröffentlicht: (2024)
A Comprehensive Survey of Contamination Detection Methods in Large Language Models
von: Ravaut, Mathieu, et al.
Veröffentlicht: (2024)
von: Ravaut, Mathieu, et al.
Veröffentlicht: (2024)
Transferable and Efficient Non-Factual Content Detection via Probe Training with Offline Consistency Checking
von: Zhang, Xiaokang, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaokang, et al.
Veröffentlicht: (2024)
"AGI" team at SHROOM-CAP: Data-Centric Approach to Multilingual Hallucination Detection using XLM-RoBERTa
von: Rathva, Harsh, et al.
Veröffentlicht: (2025)
von: Rathva, Harsh, et al.
Veröffentlicht: (2025)
MAGE: Machine-generated Text Detection in the Wild
von: Li, Yafu, et al.
Veröffentlicht: (2023)
von: Li, Yafu, et al.
Veröffentlicht: (2023)
CAP-LLM: Context-Augmented Personalized Large Language Models for News Headline Generation
von: Wilson, Raymond, et al.
Veröffentlicht: (2025)
von: Wilson, Raymond, et al.
Veröffentlicht: (2025)
DCR: Quantifying Data Contamination in LLMs Evaluation
von: Xu, Cheng, et al.
Veröffentlicht: (2025)
von: Xu, Cheng, et al.
Veröffentlicht: (2025)
Collapsed Language Models Promote Fairness
von: Xu, Jingxuan, et al.
Veröffentlicht: (2024)
von: Xu, Jingxuan, et al.
Veröffentlicht: (2024)
CreativeBench: Benchmarking and Enhancing Machine Creativity via Self-Evolving Challenges
von: Wang, Zi-Han, et al.
Veröffentlicht: (2026)
von: Wang, Zi-Han, et al.
Veröffentlicht: (2026)
Investigating Data Contamination in Modern Benchmarks for Large Language Models
von: Deng, Chunyuan, et al.
Veröffentlicht: (2023)
von: Deng, Chunyuan, et al.
Veröffentlicht: (2023)
A Logically Consistent Chain-of-Thought Approach for Stance Detection
von: Zhang, Bowen, et al.
Veröffentlicht: (2023)
von: Zhang, Bowen, et al.
Veröffentlicht: (2023)
Personality Alignment of Large Language Models
von: Zhu, Minjun, et al.
Veröffentlicht: (2024)
von: Zhu, Minjun, et al.
Veröffentlicht: (2024)
A Taxonomy for Data Contamination in Large Language Models
von: Palavalli, Medha, et al.
Veröffentlicht: (2024)
von: Palavalli, Medha, et al.
Veröffentlicht: (2024)
CAP: Evaluation of Persuasive and Creative Image Generation
von: Aghazadeh, Aysan, et al.
Veröffentlicht: (2024)
von: Aghazadeh, Aysan, et al.
Veröffentlicht: (2024)
DICE: Detecting In-distribution Contamination in LLM's Fine-tuning Phase for Math Reasoning
von: Tu, Shangqing, et al.
Veröffentlicht: (2024)
von: Tu, Shangqing, et al.
Veröffentlicht: (2024)
AutoCAP: Towards Automatic Cross-lingual Alignment Planning for Zero-shot Chain-of-Thought
von: Zhang, Yongheng, et al.
Veröffentlicht: (2024)
von: Zhang, Yongheng, et al.
Veröffentlicht: (2024)
An Open Source Data Contamination Report for Large Language Models
von: Li, Yucheng, et al.
Veröffentlicht: (2023)
von: Li, Yucheng, et al.
Veröffentlicht: (2023)
Detecting Data Contamination from Reinforcement Learning Post-training for Large Language Models
von: Tao, Yongding, et al.
Veröffentlicht: (2025)
von: Tao, Yongding, et al.
Veröffentlicht: (2025)
LaF-GRPO: In-Situ Navigation Instruction Generation for the Visually Impaired via GRPO with LLM-as-Follower Reward
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
Benchmark Data Contamination of Large Language Models: A Survey
von: Xu, Cheng, et al.
Veröffentlicht: (2024)
von: Xu, Cheng, et al.
Veröffentlicht: (2024)
Data Contamination in Neural Hieroglyphic Translation: A Reproducibility Study
von: Toutou, Ammar, et al.
Veröffentlicht: (2026)
von: Toutou, Ammar, et al.
Veröffentlicht: (2026)
Evading Data Contamination Detection for Language Models is (too) Easy
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
Towards Data Contamination Detection for Modern Large Language Models: Limitations, Inconsistencies, and Oracle Challenges
von: Samuel, Vinay, et al.
Veröffentlicht: (2024)
von: Samuel, Vinay, et al.
Veröffentlicht: (2024)
ConStat: Performance-Based Contamination Detection in Large Language Models
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
Data Contamination Can Cross Language Barriers
von: Yao, Feng, et al.
Veröffentlicht: (2024)
von: Yao, Feng, et al.
Veröffentlicht: (2024)
Quantifying Data Contamination in Psychometric Evaluations of LLMs
von: Han, Jongwook, et al.
Veröffentlicht: (2025)
von: Han, Jongwook, et al.
Veröffentlicht: (2025)
AdEval: Alignment-based Dynamic Evaluation to Mitigate Data Contamination in Large Language Models
von: Fan, Yang
Veröffentlicht: (2025)
von: Fan, Yang
Veröffentlicht: (2025)
Explainable Fake News Detection With Large Language Model via Defense Among Competing Wisdom
von: Wang, Bo, et al.
Veröffentlicht: (2024)
von: Wang, Bo, et al.
Veröffentlicht: (2024)
Data Contamination Report from the 2024 CONDA Shared Task
von: Sainz, Oscar, et al.
Veröffentlicht: (2024)
von: Sainz, Oscar, et al.
Veröffentlicht: (2024)
No Attacker Needed: Unintentional Cross-User Contamination in Shared-State LLM Agents
von: Yang, Tiankai, et al.
Veröffentlicht: (2026)
von: Yang, Tiankai, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability Curvature
von: Bao, Guangsheng, et al.
Veröffentlicht: (2023) -
LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories
von: He, Zirui, et al.
Veröffentlicht: (2025) -
Proposal Report for the 2nd SciCAP Competition 2024
von: Li, Pengpeng, et al.
Veröffentlicht: (2024) -
Unveiling the Spectrum of Data Contamination in Language Models: A Survey from Detection to Remediation
von: Deng, Chunyuan, et al.
Veröffentlicht: (2024) -
A Survey on Data Contamination for Large Language Models
von: Cheng, Yuxing, et al.
Veröffentlicht: (2025)