Improving Score Reliability of Multiple Choice Benchmarks with Consistency Evaluation and Altered Answer Choices

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cavalin, Paulo, Sanctos, Cassia, Grave, Marcelo, Pinhanez, Claudio, Primerano, Yago
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917107830095872
author Cavalin, Paulo
Sanctos, Cassia
Grave, Marcelo
Pinhanez, Claudio
Primerano, Yago
author_facet Cavalin, Paulo
Sanctos, Cassia
Grave, Marcelo
Pinhanez, Claudio
Primerano, Yago
contents In this work we present the Consistency-Rebalanced Accuracy (CoRA) metric, improving the reliability of Large Language Model (LLM) scores computed on multiple choice (MC) benchmarks. Our metric explores the response consistency of the LLMs, taking advantage of synthetically-generated questions with altered answer choices. With two intermediate scores, i.e. Bare-Minimum-Consistency Accuracy (BMCA) and Consistency Index (CI), CoRA is computed by adjusting the multiple-choice question answering (MCQA) scores to better reflect the level of consistency of the LLM. We present evaluations in different benchmarks using diverse LLMs, and not only demonstrate that LLMs can present low response consistency even when they present high MCQA scores, but also that CoRA can successfully scale down the scores of inconsistent models.
format Preprint
id arxiv_https___arxiv_org_abs_2511_21860
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving Score Reliability of Multiple Choice Benchmarks with Consistency Evaluation and Altered Answer Choices
Cavalin, Paulo
Sanctos, Cassia
Grave, Marcelo
Pinhanez, Claudio
Primerano, Yago
Computation and Language
Artificial Intelligence
In this work we present the Consistency-Rebalanced Accuracy (CoRA) metric, improving the reliability of Large Language Model (LLM) scores computed on multiple choice (MC) benchmarks. Our metric explores the response consistency of the LLMs, taking advantage of synthetically-generated questions with altered answer choices. With two intermediate scores, i.e. Bare-Minimum-Consistency Accuracy (BMCA) and Consistency Index (CI), CoRA is computed by adjusting the multiple-choice question answering (MCQA) scores to better reflect the level of consistency of the LLM. We present evaluations in different benchmarks using diverse LLMs, and not only demonstrate that LLMs can present low response consistency even when they present high MCQA scores, but also that CoRA can successfully scale down the scores of inconsistent models.
title Improving Score Reliability of Multiple Choice Benchmarks with Consistency Evaluation and Altered Answer Choices
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2511.21860