BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Balepur, Nishant, Rajasekaran, Bhavya, Oh, Jane, Xie, Michael, Desai, Atrey, Gupta, Vipul, Moore, Steven James, Choi, Eunsol, Rudinger, Rachel, Boyd-Graber, Jordan Lee
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913046834708480
author Balepur, Nishant
Rajasekaran, Bhavya
Oh, Jane
Xie, Michael
Desai, Atrey
Gupta, Vipul
Moore, Steven James
Choi, Eunsol
Rudinger, Rachel
Boyd-Graber, Jordan Lee
author_facet Balepur, Nishant
Rajasekaran, Bhavya
Oh, Jane
Xie, Michael
Desai, Atrey
Gupta, Vipul
Moore, Steven James
Choi, Eunsol
Rudinger, Rachel
Boyd-Graber, Jordan Lee
contents Multiple-choice question answering (MCQA) is standard in NLP, but benchmarks lack rigorous quality control. We present BenchMarker, an education-inspired toolkit using LLM judges to flag three common MCQ flaws: 1) contamination: items appearing exactly online; 2) shortcuts: cues in the choices that enable guessing; and 3) writing errors: structural/grammatical issues based on a 19-rule education rubric. We validate BenchMarker with human annotations, then run the tool to audit 12 benchmarks, revealing: 1) flaws persist in MCQA benchmarks, especially automatically-made and crowdsourced data - we detect 47% of TruthfulQA appears online and 100% of HellaSwag violates multiple writing rules; 2) contaminated MCQs tend to inflate accuracy, while writing errors tend to lower it and change rankings beyond random; and 3) prior benchmark repairs address their targeted issues (i.e., lowering accuracy with LLM-written distractors), but inadvertently add new flaws (i.e. implausible distractors, many correct answers). Overall, flaws in MCQs degrade NLP evaluation, but education research offers a path forward. We release BenchMarker to bridge the fields and improve MCQA benchmark design.
format Preprint
id arxiv_https___arxiv_org_abs_2602_06221
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks
Balepur, Nishant
Rajasekaran, Bhavya
Oh, Jane
Xie, Michael
Desai, Atrey
Gupta, Vipul
Moore, Steven James
Choi, Eunsol
Rudinger, Rachel
Boyd-Graber, Jordan Lee
Computation and Language
Multiple-choice question answering (MCQA) is standard in NLP, but benchmarks lack rigorous quality control. We present BenchMarker, an education-inspired toolkit using LLM judges to flag three common MCQ flaws: 1) contamination: items appearing exactly online; 2) shortcuts: cues in the choices that enable guessing; and 3) writing errors: structural/grammatical issues based on a 19-rule education rubric. We validate BenchMarker with human annotations, then run the tool to audit 12 benchmarks, revealing: 1) flaws persist in MCQA benchmarks, especially automatically-made and crowdsourced data - we detect 47% of TruthfulQA appears online and 100% of HellaSwag violates multiple writing rules; 2) contaminated MCQs tend to inflate accuracy, while writing errors tend to lower it and change rankings beyond random; and 3) prior benchmark repairs address their targeted issues (i.e., lowering accuracy with LLM-written distractors), but inadvertently add new flaws (i.e. implausible distractors, many correct answers). Overall, flaws in MCQs degrade NLP evaluation, but education research offers a path forward. We release BenchMarker to bridge the fields and improve MCQA benchmark design.
title BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks
topic Computation and Language
url https://arxiv.org/abs/2602.06221