Do LLMs Give Psychometrically Plausible Responses in Educational Assessments?

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Säuberli, Andreas, Frassinelli, Diego, Plank, Barbara
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918055076954112
author Säuberli, Andreas
Frassinelli, Diego
Plank, Barbara
author_facet Säuberli, Andreas
Frassinelli, Diego
Plank, Barbara
contents Knowing how test takers answer items in educational assessments is essential for test development, to evaluate item quality, and to improve test validity. However, this process usually requires extensive pilot studies with human participants. If large language models (LLMs) exhibit human-like response behavior to test items, this could open up the possibility of using them as pilot participants to accelerate test development. In this paper, we evaluate the human-likeness or psychometric plausibility of responses from 18 instruction-tuned LLMs with two publicly available datasets of multiple-choice test items across three subjects: reading, U.S. history, and economics. Our methodology builds on two theoretical frameworks from psychometrics which are commonly used in educational assessment, classical test theory and item response theory. The results show that while larger models are excessively confident, their response distributions can be more human-like when calibrated with temperature scaling. In addition, we find that LLMs tend to correlate better with humans in reading comprehension items compared to other subjects. However, the correlations are not very strong overall, indicating that LLMs should not be used for piloting educational assessments in a zero-shot setting.
format Preprint
id arxiv_https___arxiv_org_abs_2506_09796
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Do LLMs Give Psychometrically Plausible Responses in Educational Assessments?
Säuberli, Andreas
Frassinelli, Diego
Plank, Barbara
Computation and Language
Knowing how test takers answer items in educational assessments is essential for test development, to evaluate item quality, and to improve test validity. However, this process usually requires extensive pilot studies with human participants. If large language models (LLMs) exhibit human-like response behavior to test items, this could open up the possibility of using them as pilot participants to accelerate test development. In this paper, we evaluate the human-likeness or psychometric plausibility of responses from 18 instruction-tuned LLMs with two publicly available datasets of multiple-choice test items across three subjects: reading, U.S. history, and economics. Our methodology builds on two theoretical frameworks from psychometrics which are commonly used in educational assessment, classical test theory and item response theory. The results show that while larger models are excessively confident, their response distributions can be more human-like when calibrated with temperature scaling. In addition, we find that LLMs tend to correlate better with humans in reading comprehension items compared to other subjects. However, the correlations are not very strong overall, indicating that LLMs should not be used for piloting educational assessments in a zero-shot setting.
title Do LLMs Give Psychometrically Plausible Responses in Educational Assessments?
topic Computation and Language
url https://arxiv.org/abs/2506.09796