Metric assessment protocol in the context of answer fluctuation on MCQ tasks
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916853856600064 |
|---|---|
| author | Goliakova, Ekaterina Renard, Xavier Lesot, Marie-Jeanne Laugel, Thibault Marsala, Christophe Detyniecki, Marcin |
| author_facet | Goliakova, Ekaterina Renard, Xavier Lesot, Marie-Jeanne Laugel, Thibault Marsala, Christophe Detyniecki, Marcin |
| contents | Using multiple-choice questions (MCQs) has become a standard for assessing LLM capabilities efficiently. A variety of metrics can be employed for this task. However, previous research has not conducted a thorough assessment of them. At the same time, MCQ evaluation suffers from answer fluctuation: models produce different results given slight changes in prompts. We suggest a metric assessment protocol in which evaluation methodologies are analyzed through their connection with fluctuation rates, as well as original performance. Our results show that there is a strong link between existing metrics and the answer changing, even when computed without any additional prompt variants. A novel metric, worst accuracy, demonstrates the highest association on the protocol. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_15581 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Metric assessment protocol in the context of answer fluctuation on MCQ tasks Goliakova, Ekaterina Renard, Xavier Lesot, Marie-Jeanne Laugel, Thibault Marsala, Christophe Detyniecki, Marcin Artificial Intelligence Using multiple-choice questions (MCQs) has become a standard for assessing LLM capabilities efficiently. A variety of metrics can be employed for this task. However, previous research has not conducted a thorough assessment of them. At the same time, MCQ evaluation suffers from answer fluctuation: models produce different results given slight changes in prompts. We suggest a metric assessment protocol in which evaluation methodologies are analyzed through their connection with fluctuation rates, as well as original performance. Our results show that there is a strong link between existing metrics and the answer changing, even when computed without any additional prompt variants. A novel metric, worst accuracy, demonstrates the highest association on the protocol. |
| title | Metric assessment protocol in the context of answer fluctuation on MCQ tasks |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2507.15581 |