It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education
Fuente:
arXiv
Saved in:
| Main Authors: | Singh, Shrutika, Alyakin, Anton, Alber, Daniel Alexander, Stryker, Jaden, Tong, Ai Phuong S, Sangwon, Karl, Goff, Nicolas, de la Paz, Mathew, Hernandez-Rovira, Miguel, Park, Ki Yun, Leuthardt, Eric Claude, Oermann, Eric Karl |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MedMobile: A mobile-sized language model with clinical capabilities
by: Vishwanath, Krithik, et al.
Published: (2024)
by: Vishwanath, Krithik, et al.
Published: (2024)
BPQA Dataset: Evaluating How Well Language Models Leverage Blood Pressures to Answer Biomedical Questions
by: Hang, Chi, et al.
Published: (2025)
by: Hang, Chi, et al.
Published: (2025)
Evaluating the performance and fragility of large language models on the self-assessment for neurological surgeons
by: Vishwanath, Krithik, et al.
Published: (2025)
by: Vishwanath, Krithik, et al.
Published: (2025)
Generalist Large Language Models Outperform Clinical Tools on Medical Benchmarks
by: Vishwanath, Krithik, et al.
Published: (2025)
by: Vishwanath, Krithik, et al.
Published: (2025)
Medical large language models are easily distracted
by: Vishwanath, Krithik, et al.
Published: (2025)
by: Vishwanath, Krithik, et al.
Published: (2025)
MedG-KRP: Medical Graph Knowledge Representation Probing
by: Rosenbaum, Gabriel R., et al.
Published: (2024)
by: Rosenbaum, Gabriel R., et al.
Published: (2024)
On the Relationship Between the Choice of Representation and In-Context Learning
by: Marinescu, Ioana, et al.
Published: (2025)
by: Marinescu, Ioana, et al.
Published: (2025)
Generalist Foundation Models Are Not Clinical Enough for Hospital Operations
by: Jiang, Lavender Y., et al.
Published: (2025)
by: Jiang, Lavender Y., et al.
Published: (2025)
CNS-Obsidian: A Neurosurgical Vision-Language Model Built From Scientific Publications
by: Alyakin, Anton, et al.
Published: (2025)
by: Alyakin, Anton, et al.
Published: (2025)
"All You Need" is Not All You Need for a Paper Title: On the Origins of a Scientific Meme
by: Alyakin, Anton
Published: (2025)
by: Alyakin, Anton
Published: (2025)
Option-ID Based Elimination For Multiple Choice Questions
by: Zhu, Zhenhao, et al.
Published: (2025)
by: Zhu, Zhenhao, et al.
Published: (2025)
Questioning the Entrepreneurial State Status-quo, Pitfalls, and the Need for Credible Innovation Policy
by: Karl Wennberg
by: Karl Wennberg
Non‐invasive non‐pharmacological therapies for chronic pain: A commentary on Ikarashi et al.
by: Nabi Rustamov, et al.
Published: (2024)
by: Nabi Rustamov, et al.
Published: (2024)
Large Language Models Predict Functional Outcomes after Acute Ischemic Stroke
by: Kapoor, Anjali K., et al.
Published: (2026)
by: Kapoor, Anjali K., et al.
Published: (2026)
Refining Packing and Shuffling Strategies for Enhanced Performance in Generative Language Models
by: Chen, Yanbing, et al.
Published: (2024)
by: Chen, Yanbing, et al.
Published: (2024)
Gateformer: Advancing Multivariate Time Series Forecasting through Temporal and Variate-Wise Attention with Gated Representations
by: Lan, Yu-Hsiang, et al.
Published: (2025)
by: Lan, Yu-Hsiang, et al.
Published: (2025)
Generalization in Healthcare AI: Evaluation of a Clinical Large Language Model
by: Rahman, Salman, et al.
Published: (2024)
by: Rahman, Salman, et al.
Published: (2024)
Mitigating Easy Option Bias in Multiple-Choice Question Answering
by: Zhang, Hao, et al.
Published: (2025)
by: Zhang, Hao, et al.
Published: (2025)
Clinically Grounded Agent-based Report Evaluation: An Interpretable Metric for Radiology Report Generation
by: Dua, Radhika, et al.
Published: (2025)
by: Dua, Radhika, et al.
Published: (2025)
Testing the Solvability of Systems of Linear Inequalities
by: Goff, Leonard, et al.
Published: (2025)
by: Goff, Leonard, et al.
Published: (2025)
A spinor proof of the classification of stable minimal surfaces in $\mathbb{R}^3$
by: Stryker, Douglas
Published: (2026)
by: Stryker, Douglas
Published: (2026)
Stable 2-systole bounds in positive scalar curvature
by: Stryker, Douglas
Published: (2026)
by: Stryker, Douglas
Published: (2026)
Min-max construction of prescribed mean curvature hypersurfaces in noncompact manifolds
by: Stryker, Douglas
Published: (2024)
by: Stryker, Douglas
Published: (2024)
Lebensgeschichten alter Eltern kognitiv beeinträchtigter Menschen
by: Oermann, Lisa
Published: (2023)
by: Oermann, Lisa
Published: (2023)
Too Many or Too Few? Sampling Bounds for Topological Descriptors
by: Fasy, Brittany Terese, et al.
Published: (2025)
by: Fasy, Brittany Terese, et al.
Published: (2025)
NASA Exoplanet Exploration Program (ExEP) Science Gap List
by: Stapelfeldt, Karl, et al.
Published: (2025)
by: Stapelfeldt, Karl, et al.
Published: (2025)
NASA Exoplanet Exploration Program (ExEP) Mission Star List for the Habitable Worlds Observatory (2023)
by: Mamajek, Eric, et al.
Published: (2024)
by: Mamajek, Eric, et al.
Published: (2024)
Too Easily Fooled? Prompt Injection Breaks LLMs on Frustratingly Simple Multiple-Choice Questions
by: Guo, Xuyang, et al.
Published: (2025)
by: Guo, Xuyang, et al.
Published: (2025)
Promoting Active Citizenship Markets and Choice in Scandinavian Welfare
by: Karl Henrik Sivesind
by: Karl Henrik Sivesind
Shearing approach to gauge-invariant Trotterization
by: Stryker, Jesse R.
Published: (2021)
by: Stryker, Jesse R.
Published: (2021)
New Trends in English Education: Selected Addresses Delivered at the Conference on English Education (4th, Carnegie Institute of Technology, March 31, April 1, 2, 1966).
by: Stryker, David, Ed.
Published: (1966)
by: Stryker, David, Ed.
Published: (1966)
Race differences in support for anti‐racism policies: Endorsements of anti‐black and anti‐white stereotypes
by: Eric Silver, et al.
Published: (2024)
by: Eric Silver, et al.
Published: (2024)
Beyond racial resentment: Systemic racism beliefs and public attitudes toward criminal justice institutions and reforms
by: Eric Silver, et al.
Published: (2024)
by: Eric Silver, et al.
Published: (2024)
The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and Fixes
by: Zhu, Siqi, et al.
Published: (2026)
by: Zhu, Siqi, et al.
Published: (2026)
Choice, Psychological Ownership, and Option Valuation
by: Eugene Y. Chan
Published: (2024)
by: Eugene Y. Chan
Published: (2024)
“When Too Many Become Too Much”: Social Crowding and Its Consequences for Social Exclusion
by: Junnan Zhang, et al.
Published: (2025)
by: Junnan Zhang, et al.
Published: (2025)
The Art of Misclassification: Too Many Classes, Not Enough Points
by: Franco, Mario, et al.
Published: (2025)
by: Franco, Mario, et al.
Published: (2025)
Too Many Channels: Making Sense of Portals and Personalization.
by: Ketchell, Debra S.
Published: (2000)
by: Ketchell, Debra S.
Published: (2000)
Correcting a Nonparametric Two-sample Graph Hypothesis Test for Graphs with Different Numbers of Vertices with Applications to Connectomics
by: Alyakin, Anton A., et al.
Published: (2020)
by: Alyakin, Anton A., et al.
Published: (2020)
Paradox of De-identification: A Critique of HIPAA Safe Harbour in the Age of LLMs
by: Jiang, Lavender Y., et al.
Published: (2026)
by: Jiang, Lavender Y., et al.
Published: (2026)
Similar Items
-
MedMobile: A mobile-sized language model with clinical capabilities
by: Vishwanath, Krithik, et al.
Published: (2024) -
BPQA Dataset: Evaluating How Well Language Models Leverage Blood Pressures to Answer Biomedical Questions
by: Hang, Chi, et al.
Published: (2025) -
Evaluating the performance and fragility of large language models on the self-assessment for neurological surgeons
by: Vishwanath, Krithik, et al.
Published: (2025) -
Generalist Large Language Models Outperform Clinical Tools on Medical Benchmarks
by: Vishwanath, Krithik, et al.
Published: (2025) -
Medical large language models are easily distracted
by: Vishwanath, Krithik, et al.
Published: (2025)