Salvato in:
Dettagli Bibliografici
Autori principali: Thomas, Danielle R., Borchers, Conrad, Kakarla, Sanjit, Lin, Jionghao, Bhushan, Shambhavi, Guo, Boyuan, Gatz, Erin, Koedinger, Kenneth R.
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:https://arxiv.org/abs/2412.10267
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912154967343104
author Thomas, Danielle R.
Borchers, Conrad
Kakarla, Sanjit
Lin, Jionghao
Bhushan, Shambhavi
Guo, Boyuan
Gatz, Erin
Koedinger, Kenneth R.
author_facet Thomas, Danielle R.
Borchers, Conrad
Kakarla, Sanjit
Lin, Jionghao
Bhushan, Shambhavi
Guo, Boyuan
Gatz, Erin
Koedinger, Kenneth R.
contents The role of multiple-choice questions (MCQs) as effective learning tools has been debated in past research. While MCQs are widely used due to their ease in grading, open response questions are increasingly used for instruction, given advances in large language models (LLMs) for automated grading. This study evaluates MCQs effectiveness relative to open-response questions, both individually and in combination, on learning. These activities are embedded within six tutor lessons on advocacy. Using a posttest-only randomized control design, we compare the performance of 234 tutors (790 lesson completions) across three conditions: MCQ only, open response only, and a combination of both. We find no significant learning differences across conditions at posttest, but tutors in the MCQ condition took significantly less time to complete instruction. These findings suggest that MCQs are as effective, and more efficient, than open response tasks for learning when practice time is limited. To further enhance efficiency, we autograded open responses using GPT-4o and GPT-4-turbo. GPT models demonstrate proficiency for purposes of low-stakes assessment, though further research is needed for broader use. This study contributes a dataset of lesson log data, human annotation rubrics, and LLM prompts to promote transparency and reproducibility.
format Preprint
id arxiv_https___arxiv_org_abs_2412_10267
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Does Multiple Choice Have a Future in the Age of Generative AI? A Posttest-only RCT
Thomas, Danielle R.
Borchers, Conrad
Kakarla, Sanjit
Lin, Jionghao
Bhushan, Shambhavi
Guo, Boyuan
Gatz, Erin
Koedinger, Kenneth R.
Human-Computer Interaction
Artificial Intelligence
The role of multiple-choice questions (MCQs) as effective learning tools has been debated in past research. While MCQs are widely used due to their ease in grading, open response questions are increasingly used for instruction, given advances in large language models (LLMs) for automated grading. This study evaluates MCQs effectiveness relative to open-response questions, both individually and in combination, on learning. These activities are embedded within six tutor lessons on advocacy. Using a posttest-only randomized control design, we compare the performance of 234 tutors (790 lesson completions) across three conditions: MCQ only, open response only, and a combination of both. We find no significant learning differences across conditions at posttest, but tutors in the MCQ condition took significantly less time to complete instruction. These findings suggest that MCQs are as effective, and more efficient, than open response tasks for learning when practice time is limited. To further enhance efficiency, we autograded open responses using GPT-4o and GPT-4-turbo. GPT models demonstrate proficiency for purposes of low-stakes assessment, though further research is needed for broader use. This study contributes a dataset of lesson log data, human annotation rubrics, and LLM prompts to promote transparency and reproducibility.
title Does Multiple Choice Have a Future in the Age of Generative AI? A Posttest-only RCT
topic Human-Computer Interaction
Artificial Intelligence
url https://arxiv.org/abs/2412.10267