Tell Me Who Your Students Are: GPT Can Generate Valid Multiple-Choice Questions When Students' (Mis)Understanding Is Hinted

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shimmei, Machi, Uto, Masaki, Matsubayashi, Yuichiroh, Inui, Kentaro, Mallavarapu, Aditi, Matsuda, Noboru
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911095117053952
author Shimmei, Machi
Uto, Masaki
Matsubayashi, Yuichiroh
Inui, Kentaro
Mallavarapu, Aditi
Matsuda, Noboru
author_facet Shimmei, Machi
Uto, Masaki
Matsubayashi, Yuichiroh
Inui, Kentaro
Mallavarapu, Aditi
Matsuda, Noboru
contents The primary goal of this study is to develop and evaluate an innovative prompting technique, AnaQuest, for generating multiple-choice questions (MCQs) using a pre-trained large language model. In AnaQuest, the choice items are sentence-level assertions about complex concepts. The technique integrates formative and summative assessments. In the formative phase, students answer open-ended questions for target concepts in free text. For summative assessment, AnaQuest analyzes these responses to generate both correct and incorrect assertions. To evaluate the validity of the generated MCQs, Item Response Theory (IRT) was applied to compare item characteristics between MCQs generated by AnaQuest, a baseline ChatGPT prompt, and human-crafted items. An empirical study found that expert instructors rated MCQs generated by both AI models to be as valid as those created by human instructors. However, IRT-based analysis revealed that AnaQuest-generated questions - particularly those with incorrect assertions (foils) - more closely resembled human-crafted items in terms of difficulty and discrimination than those produced by ChatGPT.
format Preprint
id arxiv_https___arxiv_org_abs_2505_05815
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Tell Me Who Your Students Are: GPT Can Generate Valid Multiple-Choice Questions When Students' (Mis)Understanding Is Hinted
Shimmei, Machi
Uto, Masaki
Matsubayashi, Yuichiroh
Inui, Kentaro
Mallavarapu, Aditi
Matsuda, Noboru
Computation and Language
The primary goal of this study is to develop and evaluate an innovative prompting technique, AnaQuest, for generating multiple-choice questions (MCQs) using a pre-trained large language model. In AnaQuest, the choice items are sentence-level assertions about complex concepts. The technique integrates formative and summative assessments. In the formative phase, students answer open-ended questions for target concepts in free text. For summative assessment, AnaQuest analyzes these responses to generate both correct and incorrect assertions. To evaluate the validity of the generated MCQs, Item Response Theory (IRT) was applied to compare item characteristics between MCQs generated by AnaQuest, a baseline ChatGPT prompt, and human-crafted items. An empirical study found that expert instructors rated MCQs generated by both AI models to be as valid as those created by human instructors. However, IRT-based analysis revealed that AnaQuest-generated questions - particularly those with incorrect assertions (foils) - more closely resembled human-crafted items in terms of difficulty and discrimination than those produced by ChatGPT.
title Tell Me Who Your Students Are: GPT Can Generate Valid Multiple-Choice Questions When Students' (Mis)Understanding Is Hinted
topic Computation and Language
url https://arxiv.org/abs/2505.05815