Predictions from language models for multiple-choice tasks are not robust under variation of scoring methods

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tsvilodub, Polina, Wang, Hening, Grosch, Sharon, Franke, Michael
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914699568742400
author Tsvilodub, Polina
Wang, Hening
Grosch, Sharon
Franke, Michael
author_facet Tsvilodub, Polina
Wang, Hening
Grosch, Sharon
Franke, Michael
contents This paper systematically compares different methods of deriving item-level predictions of language models for multiple-choice tasks. It compares scoring methods for answer options based on free generation of responses, various probability-based scores, a Likert-scale style rating method, and embedding similarity. In a case study on pragmatic language interpretation, we find that LLM predictions are not robust under variation of method choice, both within a single LLM and across different LLMs. As this variability entails pronounced researcher degrees of freedom in reporting results, knowledge of the variability is crucial to secure robustness of results and research integrity.
format Preprint
id arxiv_https___arxiv_org_abs_2403_00998
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Predictions from language models for multiple-choice tasks are not robust under variation of scoring methods
Tsvilodub, Polina
Wang, Hening
Grosch, Sharon
Franke, Michael
Computation and Language
This paper systematically compares different methods of deriving item-level predictions of language models for multiple-choice tasks. It compares scoring methods for answer options based on free generation of responses, various probability-based scores, a Likert-scale style rating method, and embedding similarity. In a case study on pragmatic language interpretation, we find that LLM predictions are not robust under variation of method choice, both within a single LLM and across different LLMs. As this variability entails pronounced researcher degrees of freedom in reporting results, knowledge of the variability is crucial to secure robustness of results and research integrity.
title Predictions from language models for multiple-choice tasks are not robust under variation of scoring methods
topic Computation and Language
url https://arxiv.org/abs/2403.00998