WiCkeD: A Simple Method to Make Multiple Choice Benchmarks More Challenging

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Elhady, Ahmed, Agirre, Eneko, Artetxe, Mikel
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912245578989568
author Elhady, Ahmed
Agirre, Eneko
Artetxe, Mikel
author_facet Elhady, Ahmed
Agirre, Eneko
Artetxe, Mikel
contents We introduce WiCkeD, a simple method to increase the complexity of existing multiple-choice benchmarks by randomly replacing a choice with "None of the above", a method often used in educational tests. We show that WiCkeD can be automatically applied to any existing benchmark, making it more challenging. We apply WiCkeD to 6 popular benchmarks and use it to evaluate 18 open-weight LLMs. The performance of the models drops 12.1 points on average with respect to the original versions of the datasets. When using chain-of-thought on 3 MMLU datasets, the performance drop for the WiCkeD variant is similar to the one observed when using the LLMs directly, showing that WiCkeD is also challenging for models with enhanced reasoning abilities. WiCkeD also uncovers that some models are more sensitive to the extra reasoning required, providing additional information with respect to the original benchmarks. We relase our code and data at https://github.com/ahmedselhady/wicked-benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2502_18316
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WiCkeD: A Simple Method to Make Multiple Choice Benchmarks More Challenging
Elhady, Ahmed
Agirre, Eneko
Artetxe, Mikel
Computation and Language
We introduce WiCkeD, a simple method to increase the complexity of existing multiple-choice benchmarks by randomly replacing a choice with "None of the above", a method often used in educational tests. We show that WiCkeD can be automatically applied to any existing benchmark, making it more challenging. We apply WiCkeD to 6 popular benchmarks and use it to evaluate 18 open-weight LLMs. The performance of the models drops 12.1 points on average with respect to the original versions of the datasets. When using chain-of-thought on 3 MMLU datasets, the performance drop for the WiCkeD variant is similar to the one observed when using the LLMs directly, showing that WiCkeD is also challenging for models with enhanced reasoning abilities. WiCkeD also uncovers that some models are more sensitive to the extra reasoning required, providing additional information with respect to the original benchmarks. We relase our code and data at https://github.com/ahmedselhady/wicked-benchmarks.
title WiCkeD: A Simple Method to Make Multiple Choice Benchmarks More Challenging
topic Computation and Language
url https://arxiv.org/abs/2502.18316