Quantifying Data Contamination in Psychometric Evaluations of LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Han, Jongwook, Song, Woojung, Lee, Jonggeun, Jo, Yohan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Don't Adapt Small Language Models for Tools; Adapt Tool Schemas to the Models
por: Lee, Jonggeun, et al.
Publicado: (2025)
por: Lee, Jonggeun, et al.
Publicado: (2025)
Value Portrait: Assessing Language Models' Values through Psychometrically and Ecologically Valid Items
por: Han, Jongwook, et al.
Publicado: (2025)
por: Han, Jongwook, et al.
Publicado: (2025)
Human Psychometric Questionnaires Mischaracterize LLM Behavior
por: Song, Woojung, et al.
Publicado: (2025)
por: Song, Woojung, et al.
Publicado: (2025)
Psychometric Item Validation Using Virtual Respondents with Trait-Response Mediators
por: Lim, Sungjib, et al.
Publicado: (2025)
por: Lim, Sungjib, et al.
Publicado: (2025)
SpeakerSleuth: Can Large Audio-Language Models Judge Speaker Consistency across Multi-turn Dialogues?
por: Lee, Jonggeun, et al.
Publicado: (2026)
por: Lee, Jonggeun, et al.
Publicado: (2026)
SpokenUS: A Spoken User Simulator for Task-Oriented Dialogue
por: Lee, Jonggeun, et al.
Publicado: (2026)
por: Lee, Jonggeun, et al.
Publicado: (2026)
Quantifying the Effect of Test Set Contamination on Generative Evaluations
por: Schaeffer, Rylan, et al.
Publicado: (2026)
por: Schaeffer, Rylan, et al.
Publicado: (2026)
Generating Plausible Distractors for Multiple-Choice Questions via Student Choice Prediction
por: Lee, Yooseop, et al.
Publicado: (2025)
por: Lee, Yooseop, et al.
Publicado: (2025)
Mechanism Shift During Post-training from Autoregressive to Masked Diffusion Language Models
por: Kong, Injin, et al.
Publicado: (2026)
por: Kong, Injin, et al.
Publicado: (2026)
Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models
por: Riddell, Martin, et al.
Publicado: (2024)
por: Riddell, Martin, et al.
Publicado: (2024)
PVP: An Image Dataset for Personalized Visual Persuasion with Persuasion Strategies, Viewer Characteristics, and Persuasiveness Ratings
por: Kim, Junseo, et al.
Publicado: (2025)
por: Kim, Junseo, et al.
Publicado: (2025)
Sparse and Dense Retrievers Learn Better Together: Joint Sparse-Dense Optimization for Text-Image Retrieval
por: Song, Jonghyun, et al.
Publicado: (2025)
por: Song, Jonghyun, et al.
Publicado: (2025)
DCR: Quantifying Data Contamination in LLMs Evaluation
por: Xu, Cheng, et al.
Publicado: (2025)
por: Xu, Cheng, et al.
Publicado: (2025)
Format as a Prior: Quantifying and Analyzing Bias in LLMs for Heterogeneous Data
por: Liu, Jiacheng, et al.
Publicado: (2025)
por: Liu, Jiacheng, et al.
Publicado: (2025)
KL for a KL: On-Policy Distillation with Control Variate Baseline
por: Oh, Minjae, et al.
Publicado: (2026)
por: Oh, Minjae, et al.
Publicado: (2026)
Non-Collaborative User Simulators for Tool Agents
por: Shim, Jeonghoon, et al.
Publicado: (2025)
por: Shim, Jeonghoon, et al.
Publicado: (2025)
Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation
por: Mahdavi, Sadegh, et al.
Publicado: (2025)
por: Mahdavi, Sadegh, et al.
Publicado: (2025)
SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents
por: Seo, Gyuhyeon, et al.
Publicado: (2025)
por: Seo, Gyuhyeon, et al.
Publicado: (2025)
Scaling Bidirectional Spans and Span Violations in Attention Mechanism
por: Kim, Jongwook, et al.
Publicado: (2025)
por: Kim, Jongwook, et al.
Publicado: (2025)
Investigating Data Contamination for Pre-training Language Models
por: Jiang, Minhao, et al.
Publicado: (2024)
por: Jiang, Minhao, et al.
Publicado: (2024)
Data Contamination Report from the 2024 CONDA Shared Task
por: Sainz, Oscar, et al.
Publicado: (2024)
por: Sainz, Oscar, et al.
Publicado: (2024)
Time Travel in LLMs: Tracing Data Contamination in Large Language Models
por: Golchin, Shahriar, et al.
Publicado: (2023)
por: Golchin, Shahriar, et al.
Publicado: (2023)
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence
por: Choi, Hyeong Kyu, et al.
Publicado: (2025)
por: Choi, Hyeong Kyu, et al.
Publicado: (2025)
Data Contamination Quiz: A Tool to Detect and Estimate Contamination in Large Language Models
por: Golchin, Shahriar, et al.
Publicado: (2023)
por: Golchin, Shahriar, et al.
Publicado: (2023)
Investigating the Impact of Data Contamination of Large Language Models in Text-to-SQL Translation
por: Ranaldi, Federico, et al.
Publicado: (2024)
por: Ranaldi, Federico, et al.
Publicado: (2024)
Rethinking LLM Evaluation: Can We Evaluate LLMs with 200x Less Data?
por: Wang, Shaobo, et al.
Publicado: (2025)
por: Wang, Shaobo, et al.
Publicado: (2025)
Can Performant LLMs Be Ethical? Quantifying the Impact of Web Crawling Opt-Outs
por: Fan, Dongyang, et al.
Publicado: (2025)
por: Fan, Dongyang, et al.
Publicado: (2025)
Interactive and Expressive Code-Augmented Planning with Large Language Models
por: Liu, Anthony Z., et al.
Publicado: (2024)
por: Liu, Anthony Z., et al.
Publicado: (2024)
Quantifying the Capabilities of LLMs across Scale and Precision
por: Badshah, Sher, et al.
Publicado: (2024)
por: Badshah, Sher, et al.
Publicado: (2024)
Recent Advances in Large Langauge Model Benchmarks against Data Contamination: From Static to Dynamic Evaluation
por: Chen, Simin, et al.
Publicado: (2025)
por: Chen, Simin, et al.
Publicado: (2025)
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
por: Jain, Naman, et al.
Publicado: (2024)
por: Jain, Naman, et al.
Publicado: (2024)
The Emperor's New Clothes in Benchmarking? A Rigorous Examination of Mitigation Strategies for LLM Benchmark Data Contamination
por: Sun, Yifan, et al.
Publicado: (2025)
por: Sun, Yifan, et al.
Publicado: (2025)
BeyondBench: Contamination-Resistant Evaluation of Reasoning in Language Models
por: Srivastava, Gaurav, et al.
Publicado: (2025)
por: Srivastava, Gaurav, et al.
Publicado: (2025)
Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States
por: Choi, Yunho, et al.
Publicado: (2026)
por: Choi, Yunho, et al.
Publicado: (2026)
MPIB: A Benchmark for Medical Prompt Injection Attacks and Clinical Safety in LLMs
por: Lee, Junhyeok, et al.
Publicado: (2026)
por: Lee, Junhyeok, et al.
Publicado: (2026)
AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge
por: Wu, Xiaobao, et al.
Publicado: (2024)
por: Wu, Xiaobao, et al.
Publicado: (2024)
XFacta: Contemporary, Real-World Dataset and Evaluation for Multimodal Misinformation Detection with Multimodal LLMs
por: Xiao, Yuzhuo, et al.
Publicado: (2025)
por: Xiao, Yuzhuo, et al.
Publicado: (2025)
How Much Can We Forget about Data Contamination?
por: Bordt, Sebastian, et al.
Publicado: (2024)
por: Bordt, Sebastian, et al.
Publicado: (2024)
Relative Bias: A Comparative Framework for Quantifying Bias in LLMs
por: Arbabi, Alireza, et al.
Publicado: (2025)
por: Arbabi, Alireza, et al.
Publicado: (2025)
Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination
por: Wu, Mingqi, et al.
Publicado: (2025)
por: Wu, Mingqi, et al.
Publicado: (2025)
Ejemplares similares
-
Don't Adapt Small Language Models for Tools; Adapt Tool Schemas to the Models
por: Lee, Jonggeun, et al.
Publicado: (2025) -
Value Portrait: Assessing Language Models' Values through Psychometrically and Ecologically Valid Items
por: Han, Jongwook, et al.
Publicado: (2025) -
Human Psychometric Questionnaires Mischaracterize LLM Behavior
por: Song, Woojung, et al.
Publicado: (2025) -
Psychometric Item Validation Using Virtual Respondents with Trait-Response Mediators
por: Lim, Sungjib, et al.
Publicado: (2025) -
SpeakerSleuth: Can Large Audio-Language Models Judge Speaker Consistency across Multi-turn Dialogues?
por: Lee, Jonggeun, et al.
Publicado: (2026)