Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Palta, Shramay, Balepur, Nishant, Rankel, Peter, Wiegreffe, Sarah, Carpuat, Marine, Rudinger, Rachel
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917802382721024
author Palta, Shramay
Balepur, Nishant
Rankel, Peter
Wiegreffe, Sarah
Carpuat, Marine
Rudinger, Rachel
author_facet Palta, Shramay
Balepur, Nishant
Rankel, Peter
Wiegreffe, Sarah
Carpuat, Marine
Rudinger, Rachel
contents Questions involving commonsense reasoning about everyday situations often admit many $\textit{possible}$ or $\textit{plausible}$ answers. In contrast, multiple-choice question (MCQ) benchmarks for commonsense reasoning require a hard selection of a single correct answer, which, in principle, should represent the $\textit{most}$ plausible answer choice. On $250$ MCQ items sampled from two commonsense reasoning benchmarks, we collect $5,000$ independent plausibility judgments on answer choices. We find that for over 20% of the sampled MCQs, the answer choice rated most plausible does not match the benchmark gold answers; upon manual inspection, we confirm that this subset exhibits higher rates of problems like ambiguity or semantic mismatch between question and answer choices. Experiments with LLMs reveal low accuracy and high variation in performance on the subset, suggesting our plausibility criterion may be helpful in identifying more reliable benchmark items for commonsense evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2410_10854
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning
Palta, Shramay
Balepur, Nishant
Rankel, Peter
Wiegreffe, Sarah
Carpuat, Marine
Rudinger, Rachel
Computation and Language
Artificial Intelligence
Questions involving commonsense reasoning about everyday situations often admit many $\textit{possible}$ or $\textit{plausible}$ answers. In contrast, multiple-choice question (MCQ) benchmarks for commonsense reasoning require a hard selection of a single correct answer, which, in principle, should represent the $\textit{most}$ plausible answer choice. On $250$ MCQ items sampled from two commonsense reasoning benchmarks, we collect $5,000$ independent plausibility judgments on answer choices. We find that for over 20% of the sampled MCQs, the answer choice rated most plausible does not match the benchmark gold answers; upon manual inspection, we confirm that this subset exhibits higher rates of problems like ambiguity or semantic mismatch between question and answer choices. Experiments with LLMs reveal low accuracy and high variation in performance on the subset, suggesting our plausibility criterion may be helpful in identifying more reliable benchmark items for commonsense evaluation.
title Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2410.10854