Fantastic Bugs and Where to Find Them in AI Benchmarks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Truong, Sang, Tu, Yuheng, Hardy, Michael, Reuel, Anka, Tang, Zeyu, Burapacheep, Jirayu, Perera, Jonathan, Uwakwe, Chibuike, Domingue, Ben, Haber, Nick, Koyejo, Sanmi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908668469968896
author Truong, Sang
Tu, Yuheng
Hardy, Michael
Reuel, Anka
Tang, Zeyu
Burapacheep, Jirayu
Perera, Jonathan
Uwakwe, Chibuike
Domingue, Ben
Haber, Nick
Koyejo, Sanmi
author_facet Truong, Sang
Tu, Yuheng
Hardy, Michael
Reuel, Anka
Tang, Zeyu
Burapacheep, Jirayu
Perera, Jonathan
Uwakwe, Chibuike
Domingue, Ben
Haber, Nick
Koyejo, Sanmi
contents Benchmarks are pivotal in driving AI progress, and invalid benchmark questions frequently undermine their reliability. Manually identifying and correcting errors among thousands of benchmark questions is not only infeasible but also a critical bottleneck for reliable evaluation. In this work, we introduce a framework for systematic benchmark revision that leverages statistical analysis of response patterns to flag potentially invalid questions for further expert review. Our approach builds on a core assumption commonly used in AI evaluations that the mean score sufficiently summarizes model performance. This implies a unidimensional latent construct underlying the measurement experiment, yielding expected ranges for various statistics for each item. When empirically estimated values for these statistics fall outside the expected range for an item, the item is more likely to be problematic. Across nine widely used benchmarks, our method guides expert review to identify problematic questions with up to 84\% precision. In addition, we introduce an LLM-judge first pass to review questions, further reducing human effort. Together, these components provide an efficient and scalable framework for systematic benchmark revision.
format Preprint
id arxiv_https___arxiv_org_abs_2511_16842
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fantastic Bugs and Where to Find Them in AI Benchmarks
Truong, Sang
Tu, Yuheng
Hardy, Michael
Reuel, Anka
Tang, Zeyu
Burapacheep, Jirayu
Perera, Jonathan
Uwakwe, Chibuike
Domingue, Ben
Haber, Nick
Koyejo, Sanmi
Artificial Intelligence
Computation and Language
Machine Learning
Benchmarks are pivotal in driving AI progress, and invalid benchmark questions frequently undermine their reliability. Manually identifying and correcting errors among thousands of benchmark questions is not only infeasible but also a critical bottleneck for reliable evaluation. In this work, we introduce a framework for systematic benchmark revision that leverages statistical analysis of response patterns to flag potentially invalid questions for further expert review. Our approach builds on a core assumption commonly used in AI evaluations that the mean score sufficiently summarizes model performance. This implies a unidimensional latent construct underlying the measurement experiment, yielding expected ranges for various statistics for each item. When empirically estimated values for these statistics fall outside the expected range for an item, the item is more likely to be problematic. Across nine widely used benchmarks, our method guides expert review to identify problematic questions with up to 84\% precision. In addition, we introduce an LLM-judge first pass to review questions, further reducing human effort. Together, these components provide an efficient and scalable framework for systematic benchmark revision.
title Fantastic Bugs and Where to Find Them in AI Benchmarks
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2511.16842