It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916655064416256 |
|---|---|
| author | Singh, Shrutika Alyakin, Anton Alber, Daniel Alexander Stryker, Jaden Tong, Ai Phuong S Sangwon, Karl Goff, Nicolas de la Paz, Mathew Hernandez-Rovira, Miguel Park, Ki Yun Leuthardt, Eric Claude Oermann, Eric Karl |
| author_facet | Singh, Shrutika Alyakin, Anton Alber, Daniel Alexander Stryker, Jaden Tong, Ai Phuong S Sangwon, Karl Goff, Nicolas de la Paz, Mathew Hernandez-Rovira, Miguel Park, Ki Yun Leuthardt, Eric Claude Oermann, Eric Karl |
| contents | The performance of Large Language Models (LLMs) on multiple-choice question (MCQ) benchmarks is frequently cited as proof of their medical capabilities. We hypothesized that LLM performance on medical MCQs may in part be illusory and driven by factors beyond medical content knowledge and reasoning capabilities. To assess this, we created a novel benchmark of free-response questions with paired MCQs (FreeMedQA). Using this benchmark, we evaluated three state-of-the-art LLMs (GPT-4o, GPT-3.5, and LLama-3-70B-instruct) and found an average absolute deterioration of 39.43% in performance on free-response questions relative to multiple-choice (p = 1.3 * 10-5) which was greater than the human performance decline of 22.29%. To isolate the role of the MCQ format on performance, we performed a masking study, iteratively masking out parts of the question stem. At 100% masking, the average LLM multiple-choice performance was 6.70% greater than random chance (p = 0.002) with one LLM (GPT-4o) obtaining an accuracy of 37.34%. Notably, for all LLMs the free-response performance was near zero. Our results highlight the shortcomings in medical MCQ benchmarks for overestimating the capabilities of LLMs in medicine, and, broadly, the potential for improving both human and machine assessments using LLM-evaluated free-response questions. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_13508 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education Singh, Shrutika Alyakin, Anton Alber, Daniel Alexander Stryker, Jaden Tong, Ai Phuong S Sangwon, Karl Goff, Nicolas de la Paz, Mathew Hernandez-Rovira, Miguel Park, Ki Yun Leuthardt, Eric Claude Oermann, Eric Karl Computation and Language Artificial Intelligence Computers and Society The performance of Large Language Models (LLMs) on multiple-choice question (MCQ) benchmarks is frequently cited as proof of their medical capabilities. We hypothesized that LLM performance on medical MCQs may in part be illusory and driven by factors beyond medical content knowledge and reasoning capabilities. To assess this, we created a novel benchmark of free-response questions with paired MCQs (FreeMedQA). Using this benchmark, we evaluated three state-of-the-art LLMs (GPT-4o, GPT-3.5, and LLama-3-70B-instruct) and found an average absolute deterioration of 39.43% in performance on free-response questions relative to multiple-choice (p = 1.3 * 10-5) which was greater than the human performance decline of 22.29%. To isolate the role of the MCQ format on performance, we performed a masking study, iteratively masking out parts of the question stem. At 100% masking, the average LLM multiple-choice performance was 6.70% greater than random chance (p = 0.002) with one LLM (GPT-4o) obtaining an accuracy of 37.34%. Notably, for all LLMs the free-response performance was near zero. Our results highlight the shortcomings in medical MCQ benchmarks for overestimating the capabilities of LLMs in medicine, and, broadly, the potential for improving both human and machine assessments using LLM-evaluated free-response questions. |
| title | It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education |
| topic | Computation and Language Artificial Intelligence Computers and Society |
| url | https://arxiv.org/abs/2503.13508 |