Metric assessment protocol in the context of answer fluctuation on MCQ tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Goliakova, Ekaterina, Renard, Xavier, Lesot, Marie-Jeanne, Laugel, Thibault, Marsala, Christophe, Detyniecki, Marcin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916853856600064
author Goliakova, Ekaterina
Renard, Xavier
Lesot, Marie-Jeanne
Laugel, Thibault
Marsala, Christophe
Detyniecki, Marcin
author_facet Goliakova, Ekaterina
Renard, Xavier
Lesot, Marie-Jeanne
Laugel, Thibault
Marsala, Christophe
Detyniecki, Marcin
contents Using multiple-choice questions (MCQs) has become a standard for assessing LLM capabilities efficiently. A variety of metrics can be employed for this task. However, previous research has not conducted a thorough assessment of them. At the same time, MCQ evaluation suffers from answer fluctuation: models produce different results given slight changes in prompts. We suggest a metric assessment protocol in which evaluation methodologies are analyzed through their connection with fluctuation rates, as well as original performance. Our results show that there is a strong link between existing metrics and the answer changing, even when computed without any additional prompt variants. A novel metric, worst accuracy, demonstrates the highest association on the protocol.
format Preprint
id arxiv_https___arxiv_org_abs_2507_15581
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Metric assessment protocol in the context of answer fluctuation on MCQ tasks
Goliakova, Ekaterina
Renard, Xavier
Lesot, Marie-Jeanne
Laugel, Thibault
Marsala, Christophe
Detyniecki, Marcin
Artificial Intelligence
Using multiple-choice questions (MCQs) has become a standard for assessing LLM capabilities efficiently. A variety of metrics can be employed for this task. However, previous research has not conducted a thorough assessment of them. At the same time, MCQ evaluation suffers from answer fluctuation: models produce different results given slight changes in prompts. We suggest a metric assessment protocol in which evaluation methodologies are analyzed through their connection with fluctuation rates, as well as original performance. Our results show that there is a strong link between existing metrics and the answer changing, even when computed without any additional prompt variants. A novel metric, worst accuracy, demonstrates the highest association on the protocol.
title Metric assessment protocol in the context of answer fluctuation on MCQ tasks
topic Artificial Intelligence
url https://arxiv.org/abs/2507.15581