It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Singh, Shrutika, Alyakin, Anton, Alber, Daniel Alexander, Stryker, Jaden, Tong, Ai Phuong S, Sangwon, Karl, Goff, Nicolas, de la Paz, Mathew, Hernandez-Rovira, Miguel, Park, Ki Yun, Leuthardt, Eric Claude, Oermann, Eric Karl
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916655064416256
author Singh, Shrutika
Alyakin, Anton
Alber, Daniel Alexander
Stryker, Jaden
Tong, Ai Phuong S
Sangwon, Karl
Goff, Nicolas
de la Paz, Mathew
Hernandez-Rovira, Miguel
Park, Ki Yun
Leuthardt, Eric Claude
Oermann, Eric Karl
author_facet Singh, Shrutika
Alyakin, Anton
Alber, Daniel Alexander
Stryker, Jaden
Tong, Ai Phuong S
Sangwon, Karl
Goff, Nicolas
de la Paz, Mathew
Hernandez-Rovira, Miguel
Park, Ki Yun
Leuthardt, Eric Claude
Oermann, Eric Karl
contents The performance of Large Language Models (LLMs) on multiple-choice question (MCQ) benchmarks is frequently cited as proof of their medical capabilities. We hypothesized that LLM performance on medical MCQs may in part be illusory and driven by factors beyond medical content knowledge and reasoning capabilities. To assess this, we created a novel benchmark of free-response questions with paired MCQs (FreeMedQA). Using this benchmark, we evaluated three state-of-the-art LLMs (GPT-4o, GPT-3.5, and LLama-3-70B-instruct) and found an average absolute deterioration of 39.43% in performance on free-response questions relative to multiple-choice (p = 1.3 * 10-5) which was greater than the human performance decline of 22.29%. To isolate the role of the MCQ format on performance, we performed a masking study, iteratively masking out parts of the question stem. At 100% masking, the average LLM multiple-choice performance was 6.70% greater than random chance (p = 0.002) with one LLM (GPT-4o) obtaining an accuracy of 37.34%. Notably, for all LLMs the free-response performance was near zero. Our results highlight the shortcomings in medical MCQ benchmarks for overestimating the capabilities of LLMs in medicine, and, broadly, the potential for improving both human and machine assessments using LLM-evaluated free-response questions.
format Preprint
id arxiv_https___arxiv_org_abs_2503_13508
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education
Singh, Shrutika
Alyakin, Anton
Alber, Daniel Alexander
Stryker, Jaden
Tong, Ai Phuong S
Sangwon, Karl
Goff, Nicolas
de la Paz, Mathew
Hernandez-Rovira, Miguel
Park, Ki Yun
Leuthardt, Eric Claude
Oermann, Eric Karl
Computation and Language
Artificial Intelligence
Computers and Society
The performance of Large Language Models (LLMs) on multiple-choice question (MCQ) benchmarks is frequently cited as proof of their medical capabilities. We hypothesized that LLM performance on medical MCQs may in part be illusory and driven by factors beyond medical content knowledge and reasoning capabilities. To assess this, we created a novel benchmark of free-response questions with paired MCQs (FreeMedQA). Using this benchmark, we evaluated three state-of-the-art LLMs (GPT-4o, GPT-3.5, and LLama-3-70B-instruct) and found an average absolute deterioration of 39.43% in performance on free-response questions relative to multiple-choice (p = 1.3 * 10-5) which was greater than the human performance decline of 22.29%. To isolate the role of the MCQ format on performance, we performed a masking study, iteratively masking out parts of the question stem. At 100% masking, the average LLM multiple-choice performance was 6.70% greater than random chance (p = 0.002) with one LLM (GPT-4o) obtaining an accuracy of 37.34%. Notably, for all LLMs the free-response performance was near zero. Our results highlight the shortcomings in medical MCQ benchmarks for overestimating the capabilities of LLMs in medicine, and, broadly, the potential for improving both human and machine assessments using LLM-evaluated free-response questions.
title It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education
topic Computation and Language
Artificial Intelligence
Computers and Society
url https://arxiv.org/abs/2503.13508